Staff GPU Inference Engineer - Real-Time AI at Scale (Ontario)

Staff GPU Inference Engineer - Real-Time AI at Scale (Ontario)

05 Sep
|
Cerebras Systems
|
Ontario

05 Sep

Cerebras Systems

Ontario

Cerebras Systems is hiring a Software Engineer to productionize and optimize our GPU serving stack across vLLM, PyTorch, ROCm, and rack-scale GPU infrastructure. You will write production-grade code, improve time to first token, throughput, and capacity efficiency, and drive reliability in a hands-on role spanning application, runtime, distributed systems, and hardware layers.
The role emphasizes deep debugging, numerical correctness, and automated workflows to make the serving path robust and

#J-18808-Ljbffr

📌 Staff GPU Inference Engineer - Real-Time AI at Scale (Ontario)
🏢 Cerebras Systems
📍 Ontario

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: staff gpu inference engineer - real-time ai at scale (ontario) / ontario