31 Aug
|
Adaption
|
Winnipeg
As a Senior Inference Performance Engineer, you'll advance our AI stack's efficiency and reliability. Your hands-on approach to caching, batching, and optimization will be critical.
This role demands expertise accumulated over at least five years in machine learning systems or performance engineering. You will take ownership of significant performance levers while collaborating with engineers to ensure our inference models operate efficiently under varying conditions. Your contributions will focus on throughput and latency enhancements without sacrificing model integrity.
Key Responsibilities:
• Optimize throughput and latency using advanced caching techniques • Enhance performance by fine-tuning workloads with real-time data • Modify routing for cost-effective infrastructure use • Engage with frameworks such as vLLM and SGLang • Implement profiling systems for resource utilization insights
Requirements: • Minimum 5 years in ML systems or related fields • Comprehensive knowledge of model serving methodologies • Proficient in Python and at least one systems programming language • Familiar with GPU performance strategies including CUDA • Initiative-taking team player ready to innovate
Leverage your engineering skills to help us create adaptable, productive AI solutions. #J-18808-Ljbffr
📌 Senior Inference Performance Engineer (Winnipeg)
🏢 Adaption
📍 Winnipeg