Cerebras Systems is seeking a Software Engineer to productionize and optimize our GPU serving stack, spanning API services, vLLM, PyTorch, ROCm, and rack-scale AMD GPU infrastructure. You will write production code, establish operational practices, and drive improvements in time to first token, throughput, tail latency, and capacity efficiency.
You will debug across application, runtime, distributed systems, and hardware layers, and contribute to metrics, tooling, and reliability across the
#J-18808-Ljbffr
📌 Senior GPU Inference Engineer - Production-Scale AI (Toronto)
🏢 Cerebras
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.