27 Aug
|
Adaption
|
Toronto
Join our cutting-edge team as a Performance Engineer focusing on inference efficiency. Your expertise in managing KV-cache, batching, and quantization will be vital to improving our model serving capabilities. In this role, you'll directly influence the cost and performance of our inference stack.
With over five years in machine learning systems and performance engineering, you'll collaborate with software engineers to optimize throughput and latency, ensuring high model quality. Your work will involve hands-on tuning of various components under changing workloads and hardware settings. Key Responsibilities:
Enhance throughput and reduce tail latency through effective caching
Optimize workloads based on real production traffic data
Fine-tune routing between internal and external systems
Work with systems such as vLLM, SGLang, or TensorRT-LLM
Create profiling systems to analyze resource usage
Requirements:
Over 5 years in ML systems or performance engineering
In-depth knowledge of model serving dynamics
Proficient in Python and one systems language (C++, Rust)
Production experience with GPU performance optimization
Adaptable team player with a bold approach
Your skills in performance engineering will directly support our vision of adaptable and effective AI.
📌 Performance Engineer For Ai Inference Toronto
🏢 Adaption
📍 Toronto