27 Aug
|
Adaption
|
Ontario
Shape the future of AI as our Inference Performance Engineer. Your role will focus on optimizing performance metrics while working in a dynamic, collaborative environment.
You'll own the cost and performance aspects of our inference stack, leveraging over five years of experience in machine learning systems and inference infrastructure. You will closely partner with engineering teams to ensure productive model operations, tackling challenges like caching and quantization. Your results will significantly enhance model throughput and latency, maintaining top-tier quality.
Key Responsibilities:
• Manage KV-cache and continuous batching to boost performance
• Optimize prefill and decode workloads according to traffic
• Adjust routing based on performance metrics and costs
• Utilize serving engines like vLLM or TensorRT-LLM
• Develop measurement systems for time and resource usage
Requirements:
• 5+ years in ML or performance engineering
• Strong background in model serving and batching techniques
• Experienced in Python and a systems programming language
• Practical GPU performance experience including CUDA
• Creative problem solver with a collaborative spirit
Your passion for performance engineering can drive our mission of evolving AI technologies.
#J-18808-Ljbffr
📌 AI Inference Performance Engineer Role (Ontario)
🏢 Adaption
📍 Ontario