09 Aug
|
Mission.dev
|
Canada
09 Aug
Mission.dev
Canada
Agreement type : This is a remote full-time employment role requiring 40 hours per week, with direct hiring by the client.
Preferred candidates' locations : Canada
Our company description
Mission.dev is the next-gen staffing platform for software talent.
We help you find, evaluate, and manage top software talent (contractors or direct hires) faster, smarter, and more efficiently.
Powered by AI. Backed by real humans .
About the client An early-stage AI infrastructure startup specializing in inference optimization.
The company develops software to increase compute performance per watt on GPUs, integrating with major serving stacks to provide kernel-level power telemetry.
By utilizing a proprietary optimization engine and runtime controller, the platform helps AI teams maintain high throughput under power constraints, dynamically adjusting performance as models and hardware conditions evolve.
About the Role
As a Senior Performance Engineer and founding team member, you will architect solutions across the entire inference stack.
You will have direct ownership over optimizing kernels, shaping the serving engine, and improving orchestration to redefine large-scale model deployment.
Your work will directly impact the cost and energy efficiency of AI at scale, pushing GPU utilization toward theoretical limits.
This role offers the opportunity to translate complex research into robust, production-ready systems that influence the future of high-performance computing.
What You'll Do
Design and build high-performance inference systems for large-scale models across multi-GPU and multi-node deployments.
Develop custom kernels and runtime paths to maximize GPU utilization and memory efficiency.
Extend serving engines by implementing advanced batching, KV-cache management,
and scheduling policies.
Architect distributed inference systems including autoscaling, load balancing, and failure handling under strict latency SLOs.
Profile and analyze the full inference lifecycle to identify and resolve systemic bottlenecks.
Implement advanced techniques such as speculative decoding, tensor parallelism, and mixture-of-experts serving.
Collaborate with hardware partners to co-design software strategies that improve energy efficiency and reduce operational costs.
Contribute to internal tooling and relevant open-source projects within the inference ecosystem.
What You Bring
Extensive experience building or operating large-scale production inference or training systems.
Deep understanding of GPU architectures, including compute, memory bandwidth, and kernel launch characteristics.
Proficiency with the modern inference ecosystem and frameworks.
Robust foundation in systems and distributed systems, including concurrency, scheduling, and fault tolerance.
Expertise in debugging performance issues across layers, from kernel traces to autoscaler dynamics.
Ability to own technical problems end-to-end, from initial measurement to long-term operational health.
Nice to Haves
Contributions to major open-source inference frameworks or libraries.
Experience with Kubernetes-based infrastructure and modern observability stacks.
Knowledge of energy-aware scheduling or capacity planning for large accelerator fleets.
Familiarity with post-training pipelines and their interaction with inference infrastructure.
Compensation & Benefits
Founding-engineer equity and direct ownership.
Access to state-of-the-art compute resources and hardware.
Remote-friendly. Montreal preferred for regular in-person work with the founding team.
📌 Senior Performance Engineer - LLM Inference Optimization (Canada)
🏢 Mission.dev
📍 Canada