02 Oct
|
Atomic - Remote Jobs
|
Vancouver
02 Oct
Atomic - Remote Jobs
Vancouver
Company Overview Our client is a fast-growing applied research lab building the data layer for frontier AI. They partner with leading AI labs and enterprises to deliver two things: proprietary, expert-generated datasets, and rigorous evaluation and benchmarking. The goal is for AI systems to get better at real workflows, not just polished demos.
Founded in 2025 and backed by a $30M Series A, the company is fully remote. It works on problems such as turning real-world work into clean training signals, building evaluations for software engineering agents and finance workflows, and proving data quality by measuring actual performance lift.
Your Role This is not a "pick up tickets and wait for specs" role.
- This is a broad builder seat that combines platform engineering with evaluation and experimentation infrastructure. You'll design, build, and run systems in production that researchers and operators depend on every day.
- The company is scaling the pipelines, evaluation harnesses, and training environments that make its work repeatable, and it needs engineers who can own a system and deliver reliably.
- In your first 30 to 90 days, success means shipping at least one meaningful production improvement and becoming the go-to owner of a core system.
You'll:
- Build and maintain evaluation harnesses that measure model and agent performance on real tasks
- Improve eval reliability, coverage, and signal quality through better rubrics, task design support, and scoring
- Ship tools that let researchers and operators run experiments without reinventing the process each time
- Build APIs and backend services that power human-in-the-loop workflows, task routing, and quality checks
- Improve the pipelines that turn expert work into structured training and evaluation data
- Make systems more observable, scalable, and easier to operate through logging, metrics,
and debugging
- Write explicit, maintainable code, take part in reviews and design discussions, and document decisions so others can build on them
You Bring:
- Strong coding fundamentals in Node.js and TypeScript
- Strong coding ability in Python and/or Go
- Experience building and owning production systems such as APIs, services, and pipelines
- A solid understanding of distributed systems and engineering trade-offs
- Comfort with AWS or GCP and modern infrastructure (containers, Kubernetes)
- A track record of shipping and maintaining systems other people rely on, not only prototypes
- Solid written communication and comfort working async in a distributed team
Bonus Points:
- Experience with evaluation frameworks, experimentation platforms, or ML tooling
- Experience with data pipelines or workflow orchestration
- Experience building internal platforms for operators or research teams
- Experience in early-stage or high-ownership B2B SaaS or platform teams
What's Offered:
- Full-time, fully remote role with a LATAM focus and meaningful overlap with U.S. time zones
- $7,000–10,000 USD/month, based on experience
- Real ownership, with growth into bigger systems, deeper technical leadership, and projects core to how the company scales
- A lean, async-first team that values clear writing, sound judgment, and follow-through
- Research-adjacent engineering at the frontier of AI, alongside practical platform work
- Direct impact on the data and evaluations used by leading AI labs
Interview Process: 1️⃣ Take-home assignment covering practical engineering and how you communicate decisions 2️⃣ Application review by the team
3️⃣ Founding engineer screen, a deep dive on system design, trade-offs, and past ownership
4️⃣ Work trial on real-world work, focused on execution, quality, and collaboration
5️⃣ Offer
📌 Software Engineer | AI Training Data & Evals Lab ? (Vancouver)
🏢 Atomic - Remote Jobs
📍 Vancouver