31 Aug
|
Adaption
|
Toronto
Shape the future of AI as our Inference Performance Engineer.
Your role will focus on optimizing performance metrics while working in a dynamic, cooperative setting. You'll own the cost and performance aspects of our inference stack, leveraging over five years of experience in machine learning systems and inference infrastructure. You will closely partner with engineering teams to ensure efficient model operations, tackling challenges like caching and quantization.
Your results will significantly enhance model throughput and latency, maintaining top-tier quality. Key Responsibilities:
Manage KV-cache and continuous batching to boost performance
Optimize prefill and decode workloads according to traffic
Adjust routing based on performance metrics and costs
Utilize serving engines like vLLM or TensorRT-LLM
Develop measurement systems for time and resource usage
Requirements:
5+ years in ML or performance engineering
Solid background in model serving and batching techniques
Experienced in Python and a systems programming language
Practical GPU performance experience including CUDA
Creative problem solver with a team-oriented spirit
Your passion for performance engineering can drive our mission of evolving AI technologies.
📌 Ai Inference Performance Engineer Role Toronto
🏢 Adaption
📍 Toronto