Ai Inference Performance Engineer Role Toronto

Ai Inference Performance Engineer Role Toronto

02 Sep
|
Adaption
|
Toronto

02 Sep

Adaption

Toronto

Shape the future of AI as our Inference Performance Engineer. Your role will focus on optimizing performance metrics while working in a energetic, team-oriented environment.

You'll own the cost and performance aspects of our inference stack, leveraging over five years of experience in machine learning systems and inference infrastructure. You will closely partner with engineering teams to ensure effective model operations, tackling challenges like caching and quantization. Your results will significantly enhance model throughput and latency, maintaining top-tier quality.

Key Responsibilities
Manage KV-cache and continuous batching to boost performance
Optimize prefill and decode workloads according to traffic




Adjust routing based on performance metrics and costs
Utilize serving engines like vLLM or TensorRT-LLM
Develop measurement systems for time and resource usage

Requirements:
5+ years in ML or performance engineering
Strong background in model serving and batching techniques
Experienced in Python and a systems programming language
Practical GPU performance experience including CUDA
Creative problem solver with a collaborative spirit

Your passion for performance engineering can drive our mission of evolving AI technologies.

📌 Ai Inference Performance Engineer Role Toronto
🏢 Adaption
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai inference performance engineer role toronto / toronto