Senior. AI Engineer – LLM Research Engineer (Toronto)

Senior. AI Engineer – LLM Research Engineer (Toronto)

06 Aug
|
Chubb
|
Toronto

06 Aug

Chubb

Toronto

Senior LLM Research Engineer About Chubb Chubb is the world's largest publicly traded property and casualty insurer, with operations in 54 countries providing commercial and personal insurance, reinsurance, and life insurance to clients worldwide. We're at the forefront of AI-driven transformation in insurance. Chubb is making strategic investments in artificial intelligence to fundamentally change how we assess risk, serve customers, and operate our business.

Combining cutting-edge AI capabilities with our exceptional financial strength, comprehensive product portfolio, and global reach, we're building the future of intelligent insurance solutions. The Role

As a Senior LLM Research Engineer you'll own the loop from dataset to trained checkpoint to measured result. Post-training and inference performance sit in one role because in practice they are one loop: train, quantize, serve, measure, then feed the result into the next run. What makes it interesting is what you train against.

Underwriting, claims, and the other core areas each bring their own vocabulary and their own definition of a correct answer, across diverse language tasks with an emphasis on long-form reasoning and complex instruction following.

Ground truth largely does not exist yet: the corpora are large and were never assembled for training, so building training and evaluation sets from them is a novel problem in its own right. We are data-centric by conviction, and post-training data quality is where most of our gains come from. This is applied science aimed at direct implementation.

People on this team move between training, systems engineering, and agentic work as priorities shift. We run closer to a dense model than a mixture of experts.

Major





Responsibilities

Run post-training end-to-end: SFT, GRPO and other RL methods, and distillation

Design and build evaluation: judge rubrics, eval harnesses, regression suites, and the analysis that turns a training run into a decision

Build ground truth where none exists: construct labelled training and evaluation sets from existing corpora, including data synthesis at scale, define what a correct answer looks like per task, and validate that the sets measure what they claim to. Data quality decides whether a run is worth doing

Run distributed training on multi-node GPU clusters: parallelism strategy, sharding and offload, memory and throughput tuning, and debugging runs that fail at hour nine

Inference performance engineering: tensor and pipeline parallel configuration on vLLM, quantization, speculative decoding, and throughput tuning for both evaluation and serving What You'll Bring

Deep expertise in production Python, with strong working knowledge of the training and inference stack (PyTorch, Transformers, TRL, DeepSpeed, vLLM, or equivalents)

Hands-on post-training experience. You've run SFT and at least one RL method on a real model, and can explain what went wrong the first time

Experience with multi-node distributed training: parallelism strategy, sharding and offload,



and the failure modes that only appear at scale

Evaluation rigour, and the judgement to design measurement that survives contact with a business workplace

Fluency with agentic coding tools in your own workflow, such as Claude Code and Codex. The team is fully immersed in this way of engineering, and we expect it to be part of how you build rather than something reached for occasionally The profile can come from either direction: a machine learning engineer who has gone deep on training, or an ML researcher whose work has made it to impact stage. Either way we expect engineering discipline, meaning version control, tests, reproducibility, and code the next person can pick up Strong Preference Given To

Building labelled training and evaluation sets from unlabelled source data, including synthetic data generation

LLM-as-judge evaluation design, including inter-rater agreement and rubric validation

Inference optimization: quantization, speculative decoding, or kernel-level performance work

Domain-heavy language tasks where the correct answer is ambiguous and nuanced At Chubb we are committed to providing equal employment opportunities to all employees and applicants. It is our policy to provide equal employment opportunities to employees and applicants based on job-related qualifications and ability to perform a job. If you require an accommodation during the hiring process or upon hire, please inform Human Resources.

If a selected applicant requests accommodation during the recruitment process, Chubb will consult with the applicant in order to provide suitable accommodation that takes into account the applicant’s accessibility needs.

📌 Senior. AI Engineer – LLM Research Engineer (Toronto)
🏢 Chubb
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior. ai engineer – llm research engineer (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: senior. ai engineer – llm research engineer (toronto) / toronto