Join our cutting-edge team as a Performance Engineer focusing on inference efficiency. Your expertise in managing KV-cache, batching, and quantization will be vital to improving our model serving capabilities.
In this role, you'll directly influence the cost and performance of our inference stack. With over five years in machine learning systems and performance engineering, you'll collaborate with software engineers to optimize throughput and latency, ensuring high model quality. Your work will involve hands-on tuning of various components under changing workloads and hardware environments.
Key Responsibilities:
• Enhance throughput and reduce tail latency through effective caching
• Optimize workloads based on real production traffic data
• Fine-tune routing between internal and external systems
• Work with systems such as vLLM, SGLang, or TensorRT-LLM
• Create profiling systems to analyze resource usage
Requirements:
• Over 5 years in ML systems or performance engineering
• In-depth knowledge of model serving dynamics
• Proficient in Python and one systems language (C++, Rust)
• Production experience with GPU performance optimization
• Adaptable team player with a bold approach
Your skills in performance engineering will directly support our vision of flexible and effective AI.
#J-18808-Ljbffr
📌 Performance Engineer for AI Inference (Toronto)
🏢 Adaption
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.