· Location: Remote within the Canada
· Type: 6-month contract, W-2, opportunity to convert to full-time
· Level: Masters student or early career About Us We're a stealth-mode startup based out of Seattle, USA building an AI agent platform for healthcare SaaS — LLM-powered agents that operate real web applications, process complex clinical and claims documents, and execute multi-step back-office workflows end to end. We're a small early-stage team with paying enterprise customers, shipping fast in a domain where correctness, auditability, and compliance actually matter. Full details on the company, product, and customers shared during interviews.
The Role
We're looking for an ML Engineer Intern to work alongside the founding team on our applied-LLM systems — agent evaluation, document AI, prompt iteration, and production observability.
To be clear about the shape of the work: this is applied LLM engineering, not model training or research. You won't be running pretraining jobs or writing papers. You'll be building the evaluation harnesses and data pipelines that tell us whether our agents actually work, and shipping that into production alongside engineers who've done it before.
This is not a scoped side project that gets thrown away in August. You'll own real surface area on systems that run against live customer workflows, with a mentor and weekly design reviews. Six months, with a full-time offer possible at or before the end based on performance and mutual fit.
What You'll Do Build and extend our evaluation suite: golden datasets, adversarial and distractor cases, and deterministic accuracy benchmarks that gate CI — so model and prompt changes ship with evidence, not vibes.
Curate and label evaluation data from real workflow traces,
working with domain experts to encode what a correct decision actually looks like.
Iterate on planner prompts and tool-calling loops inside our agent orchestration layer (we use LangGraph), and measure what your changes did.
Contribute to document AI pipelines: OCR and layout-aware extraction (Docling), table extraction, and grounded document Q&A; over clinical and claims documents.
Instrument and debug production LLM behavior with tracing and observability (Langfuse), including PHI-safe trace sanitization.
Help benchmark our multi-provider model layer — latency, cost, and quality trade-offs across frontier models (Gemini, GPT, Claude families) through an OpenAI-compatible gateway.
Work with embeddings and vector search (pgvector) for agent memory and pattern reuse.
Practice HIPAA-aware engineering: PHI redaction, encryption, audit logging — compliance is a feature here, not an afterthought.
What We're Looking For Recently completed a BS/MS in CS, ML, or a related field.
Robust Python.
Working
TypeScript/Node.js, or the appetite to pick it up quickly (our stack is Node/Express/Socket.io + React).
You've built something real with LLMs beyond a notebook or a demo — a side project, research tool, or project you took to actual users. We care much more about depth on one thing than breadth across many.
Practical familiarity with tool/function calling, structured outputs, and prompt engineering.
You can explain how you'd know a change made a system better.
This is the single most important thing we screen for.
Comfort with ambiguity and early-stage pace: you scope your own work, ask questions early, and communicate clearly in writing.
Available roughly 40 hours/week for the full 6 months. We'll consider a reduced schedule for a strong candidate balancing coursework. Nice to Have Browser automation or DOM-level agent experience (Playwright/CDP, or LLM-driven UI agents).
Document AI: OCR, layout parsing (Docling, Textract), information extraction from messy real-world documents.
Healthcare, insurance, or other regulated-domain exposure (claims, prior auth, clinical records).
LangGraph/LangChain, Langfuse/Braintrust, OpenRouter, or MCP experience.
Fine-tuning, distillation, or RL/evals background.
Real-time systems: WebSockets, streaming LLM responses.
Contract Details Duration: 6 months, with a defined start and end date
Conversion: Full-time offer possible at or before the end, based on performance and mutual fit
Compensation: $4,000-$6000/month, paid as a W-2 employee
Work mode: Hybrid in the Seattle area preferred; remote within the US considered for the right candidate
Equipment: Company-provided laptop
Work authorization: You must be authorized to work in the US for the duration of the contract. F-1 students on CPT or OPT are welcome to apply — please note your status and any school approval timelines in your application.
How to Apply Send your resume or LinkedIn plus a short note about an LLM system you built — what you owned, what broke, and how you measured quality — to
[email protected]. If you don't have an LLM project yet, tell us about the hardest bug you've ever debugged and how you found it. We read every application.
📌 Machine Learning Engineer — Applied LLM / Agents (Contract) (Vancouver)
🏢 Stealth Startup
📍 Vancouver