Cohere is a security-first AI company focused on building effective foundation models and end-to-end products for enterprises. This role targets optimizing the inference stack, reducing latency, and boosting throughput for large language models.
You will work with researchers and engineers across modeling and systems teams, exploring GPU-accelerated optimizations and MoE strategies to accelerate production deployments.
#J-18808-Ljbffr
📌 Staff ML Systems Engineer - Model Efficiency & Inference (Toronto)
🏢 Cohere
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.