Join our cutting-edge team as a Performance Engineer focusing on inference efficiency. Your expertise in managing KV-cache, batching, and quantization will be vital to improving our model serving capabilities.In this role, you'll directly influence the cost and performance of our inference stack. With over five years in machine learning systems and performance engineering, you'll collaborate with software engineers to optimize throughput and latency, ensuring high model quality.
Your work will involve hands-on tuning of various components under changing workloads and hardware environments.Key Responsibilities:Enhance throughput and reduce tail latency through effective cachingOptimize workloads based on real production traffic dataFine-tune routing between internal and external systemsWork with systems such as vLLM, SGLang, or TensorRT-LLMCreate profiling systems to analyze resource usageRequirements:Over 5 years in ML systems or performance engineeringIn-depth knowledge of model serving dynamicsProficient in Python and one systems language (C++, Rust)Production experience with GPU performance optimizationAdaptable team player with a bold approachYour skills in performance engineering will directly support our vision of versatile and efficient AI.#J-18808-Ljbffr
📌 Performance Engineer For Ai Inference (Toronto)
🏢 Adaption
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.