Senior LLM Research Engineer About Chubb Chubb is the world's largest publicly traded property and casualty insurer, with operations in 54 countries providing commercial and personal insurance, reinsurance, and life insurance to clients worldwide. We're at the forefront of AI-driven transformation in insurance. Chubb is making strategic investments in artificial intelligence to fundamentally change how we assess risk, serve customers, and operate our business.
Combining cutting-edge AI capabilities with our exceptional financial strength, comprehensive product portfolio, and global reach, we're building the future of intelligent insurance solutions. The Role
As a Senior LLM Research Engineer you'll own the loop from dataset to trained checkpoint to measured result. Post-training and inference performance sit in one role because in practice they are one loop: train, quantize, serve, measure, then feed the result into the next run. What makes it interesting is what you train against.
Underwriting, claims, and the other core areas each bring their own vocabulary and their own definition of a correct answer, across diverse language tasks with an emphasis on long-form reasoning and complex instruction following.
Ground truth largely does not exist yet: the corpora are large and were never assembled for training, so building training and evaluation sets from them is a novel problem in its own right. We are data-centric by conviction, and post-training data quality is where most of our gains come from. This is applied science aimed at direct implementation.
People on this team move between training, systems engineering, and agentic work as priorities shift. We run closer to a dense model than a mixture of experts.
Major
Responsibilities
Run post-training end-to-end: SFT, GRPO and other RL methods, and distillation
Design and build evaluation: judge rubrics, eval harnesses, regression suites, and the analysis that turns a training run into a decision
Build ground truth where none exists: construct labelled training and evaluation sets from existing corpora, including data synthesis at scale, define what a correct answer looks like per task, and validate that the sets measure what they claim to. Data quality decides whether a run is worth doing
Run distributed training on multi-node GPU clusters: parallelism strategy, sharding and offload, memory and throughput tuning, and debugging runs that fail at hour nine
Inference performance engineering: tensor and pipeline parallel configuration on vLLM, quantization, speculative decoding, and throughput tuning for both evaluation and serving What You'll Bring
Deep expertise in production Python, with strong working knowledge of the training and inference stack (PyTorch, Transformers, TRL, DeepSpeed, vLLM, or equivalents)
Hands-on post-training experience. You've run SFT and at least one RL method on a real model, and can explain what went wrong the first time
Experience with multi-node distributed training: parallelism strategy, sharding and offload,
and the failure modes that only appear at scale
Evaluation rigour, and the judgement to design measurement that survives contact with a business workplace
Fluency with agentic coding tools in your own workflow, such as Claude Code and Codex. The team is fully immersed in this way of engineering, and we expect it to be part of how you build rather than something reached for occasionally The profile can come from either direction: a machine learning engineer who has gone deep on training, or an ML researcher whose work has made it to impact stage. Either way we expect engineering discipline, meaning version control, tests, reproducibility, and code the next person can pick up Strong Preference Given To
Building labelled training and evaluation sets from unlabelled source data, including synthetic data generation
LLM-as-judge evaluation design, including inter-rater agreement and rubric validation
Inference optimization: quantization, speculative decoding, or kernel-level performance work
Domain-heavy language tasks where the correct answer is ambiguous and nuanced At Chubb we are committed to providing equal employment opportunities to all employees and applicants. It is our policy to provide equal employment opportunities to employees and applicants based on job-related qualifications and ability to perform a job. If you require an accommodation during the hiring process or upon hire, please inform Human Resources.
If a selected applicant requests accommodation during the recruitment process, Chubb will consult with the applicant in order to provide suitable accommodation that takes into account the applicant’s accessibility needs.
📌 Senior. AI Engineer – LLM Research Engineer (Toronto)
🏢 Chubb
📍 Toronto