31 Aug
|
Adaption
|
Toronto
As a Senior Inference Performance Engineer, you'll advance our AI stack's efficiency and reliability. Your hands-on approach to caching, batching, and optimization will be critical.This role demands expertise accumulated over at least five years in machine learning systems or performance engineering. You will take ownership of significant performance levers while collaborating with engineers to ensure our inference models operate efficiently under varying conditions.
Your contributions will focus on throughput and latency enhancements without sacrificing model integrity.Key Responsibilities:Optimize throughput and latency using advanced caching techniquesEnhance performance by fine-tuning workloads with real-time dataModify routing for cost-effective infrastructure useEngage with frameworks such as vLLM and SGLangImplement profiling systems for resource utilization insightsRequirements:Minimum 5 years in ML systems or related fieldsComprehensive knowledge of model serving methodologiesProficient in Python and at least one systems programming languageFamiliar with GPU performance strategies including CUDAInitiative-taking team player ready to innovateLeverage your engineering skills to help us create adaptable, productive AI solutions.
📌 Senior Inference Performance Engineer (Toronto)
🏢 Adaption
📍 Toronto