Machine Learning Software Engineer (British Columbia)

Machine Learning Software Engineer (British Columbia)

09 Oct
|
EMA
|
British Columbia

09 Oct

EMA

British Columbia

Ema builds AI employees that carry out complex workflows across enterprise applications. Our ML team works on the loop that makes them better: production traces become data, data becomes training and evaluation, and better agents produce better traces. The hard part is deciding which intervention will improve behavior in the next real workflow
Harnesses and inference-time compute. Design context, tools, skills and orchestration for multi-step agents, including work across documents, slides, images, audio and video. Test where extra reasoning, search or verification earns its latency and cost. Build self-improvement loops with explicit permissions, evaluation gates and rollback
Agent post-training. Curate trajectories for SFT, optimize preferences, or run RL on real agent tasks. Investigate methods such as DPO, GRPO or DAPO where they fit; compare process and outcome supervision, shape rewards, and distill useful frontier behavior into smaller models. Measure whether gains transfer beyond the training environment
Environments and rewards. Turn enterprise workflows into reproducible training and evaluation environments: fixture tenants, simulated users who may get impatient and leave, and rewards grounded in verifiable outcomes. Find the shortcuts an agent can exploit before a training run optimizes for them
Data engines and evaluation. Mine production agent-steps for failures; build curated corpora and practical synthetic augmentation. Calibrate judges against human labels, construct behavior-level benchmarks from real workflows, and quantify data quality, performance uplift and reliability across stochastic runs
Retrieval, memory and context graphs. Connect enterprise information with user- and tenant-level learnings. Separate failures of retrieval from failures to use retrieved context; test what to retain,



update and retrieve so that past experience improves the next decision
Quality per dollar. Build and evaluate routing, ensembles, caching and small-model specialization. Measure downstream task success alongside latency and cost; a cheaper model is useful only if the complete agent still succeeds
You’ll go deep in a subset of these areas. Projects combine applied research with the engineering needed to make the result work in production
Most projects develop over roughly four to six months, with useful improvements shipping along the way. You’ll define the problem and baseline, build the data or environment needed to test it, run experiments, and own serving and integration
Follow the system through deployment, monitoring and failure analysis until it is ready for a clear engineering handoff
You’ll work with researchers and engineers across the Bay Area, Vancouver and India. We’re hiring from junior through senior levels, with project scope matched to your experience
A master’s or PhD in a relevant field, or equivalent work or research experience. Papers, substantial open-source contributions, trained models and well-documented experiments can demonstrate that depthStatistical judgment. You can size an experiment, choose meaningful baselines and held-out tests, and account for variation across tasks, seeds and repeated runs. You can distinguish a real improvement from judge bias,



data leakage or a benchmark shortcutEvidence of zero-to-one ownership. A system, model or research project you took from an ambiguous problem to a working result. We want to understand your contribution, the tradeoffs you made, and what changed when the work met real users or realistic tasksThese are project-specific strengths; post-training experience is optional, and no candidate needs the entire list. Bring a repository, paper, model or technical write-up that lets us examine how you think and what you builtDepth you can defend. Substantial work in at least one of agent/tool-use systems, post-training, reward modeling or RL environments, retrieval and memory, or evaluation design. Be ready to explain the mechanism, the alternatives you rejected and the failure modes you found. One area you can teach us beats five you’ve touchedProduction engineering judgment. You can debug across the model and system boundary, isolate a failure, and turn the result into maintainable production codeHonest measurement. You would rather retire your own approach after a clean negative result than ship an improvement that disappears under a stronger evaluationInteractive agent environments/harnesses for software engineering, web or tool use; large-scale trace analysis, data curation or synthetic generationPractical security work on prompt injection, data governance or permission boundaries for agents that can act and improve themselvesOpen-model post-training with TRL, veRL, OpenRLHF, or similar or a custom loop, especially debugging reward hacking or unstable optimizationDesigning systems to support complex, long-horizon agent work across a multitude of modalities and platformsServing with vLLM or SGLang, distillation, quantization, or multi-node GPU training

#J-18808-Ljbffr

📌 Machine Learning Software Engineer (British Columbia)
🏢 EMA
📍 British Columbia

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: machine learning software engineer (british columbia) / british columbia

Subscribe to this job alert:

Get the latest job offers by email for: machine learning software engineer (british columbia) / british columbia