Senior LLM Research EngineerAbout ChubbChubb is the world's largest publicly traded property and casualty insurer, with operations in 54 countries providing commercial and personal insurance, reinsurance, and life insurance to clients worldwide.We're at the forefront of AI-driven transformation in insurance. Chubb is making strategic investments in artificial intelligence to fundamentally change how we assess risk, serve customers, and operate our business. Combining cutting-edge AI capabilities with our exceptional financial strength, comprehensive product portfolio, and global reach, we're building the future of intelligent insurance solutions.The RoleAs a Senior LLM Research Engineer you'll own the loop from dataset to trained checkpoint to measured result. Post-training and inference performance sit in one role because in practice they are one loop: train, quantize, serve, measure, then feed the result into the next run.What makes it interesting is what you train against. Underwriting, claims, and the other core areas each bring their own vocabulary and their own definition of a correct answer, across diverse language tasks with an emphasis on long-form reasoning and complex instruction following. Ground truth largely does not exist yet: the corpora are large and were never assembled for training, so building training and evaluation sets from them is a novel problem in its own right. We are data-centric by conviction, and post-training data quality is where most of our gains come from.This is applied science aimed at direct implementation. People on this team move between training, systems engineering, and agentic work as priorities shift.
We run closer to a dense model than a mixture of experts.Major ResponsibilitiesRun post-training end-to-end: SFT, GRPO and other RL methods, and distillationDesign and build evaluation: judge rubrics, eval harnesses, regression suites, and the analysis that turns a training run into a decisionBuild ground truth where none exists: construct labelled training and evaluation sets from existing corpora, including data synthesis at scale, define what a correct answer looks like per task, and validate that the sets measure what they claim to. Data quality decides whether a run is worth doingRun distributed training on multi-node GPU clusters: parallelism strategy, sharding and offload, memory and throughput tuning, and debugging runs that fail at hour nineInference performance engineering: tensor and pipeline parallel configuration on vLLM, quantization, speculative decoding, and throughput tuning for both evaluation and servingWhat You'll BringDeep expertise in production Python, with solid working knowledge of the training and inference stack (PyTorch, Transformers, TRL, DeepSpeed, vLLM, or equivalents)Hands-on post-training experience. You've run SFT and at least one RL method on a real model, and can explain what went wrong the first timeExperience with multi-node distributed training: parallelism strategy, sharding and offload,
and the failure modes that only appear at scaleEvaluation rigour, and the judgement to design measurement that survives contact with a business environmentFluency with agentic coding tools in your own workflow, such as Claude Code and Codex. The team is fully immersed in this way of engineering, and we expect it to be part of how you build rather than something reached for occasionallyThe profile can come from either direction: a machine learning engineer who has gone deep on training, or an ML researcher whose work has made it to impact stage. Either way we expect engineering discipline, meaning version control, tests, reproducibility, and code the next person can pick upStrong Preference Given ToBuilding labelled training and evaluation sets from unlabelled source data, including synthetic data generationLLM-as-judge evaluation design, including inter-rater agreement and rubric validationInference optimization: quantization, speculative decoding, or kernel-level performance workDomain-heavy language tasks where the correct answer is ambiguous and nuancedAt Chubb we are committed to providing equal employment opportunities to all employees and applicants. It is our policy to provide equal employment opportunities to employees and applicants based on job-related qualifications and ability to perform a job. If you require an accommodation during the hiring process or upon hire, please inform Human Resources. If a selected applicant requests accommodation during the recruitment process, Chubb will consult with the applicant in order to provide suitable accommodation that takes into account the applicant's accessibility needs. #J-18808-Ljbffr
📌 Senior. Ai Engineer – Llm Research Engineer (Toronto)
🏢 Chubb
📍 Toronto