Backend Engineer — Model Inference (Ontario)

Backend Engineer — Model Inference (Ontario)

30 Sep
|
Best AI Tools Wiki
|
Ontario

30 Sep

Best AI Tools Wiki

Ontario

Cohere is hiring a Backend Engineer to optimize model inference systems. You will build low-latency serving infrastructure, implement batching strategies, and develop the backend services that power Cohere's enterprise API.
Requirements 5+ years of backend engineering experience
Strong proficiency in Go, Rust, or C++
Experience with high-throughput, low-latency systems
Knowledge of model serving and inference optimization
Experience with gRPC, REST APIs, and microservices
Nice to Have Experience with vLLM, TensorRT-LLM, or similar
Familiarity with GPU memory management
Experience with continuous batching techniques
Equity in a well-funded AI startup Premium health and dental
Remote-first culture
Home office budget
Annual learning stipend
Versatile PTO
Skills Go Rust Model Serving Backend Engineering gRPC Inference Optimization
This listing was posted on 2026-02-17 and is almost certainly closed. It is kept for reference — check the employer’s own careers page for current openings.
Vincony has all 400+ AI models in one place — compare responses, AI debate, Image/Video/Voice generator, and 20 more tools to help you learn and build with AI.

#J-18808-Ljbffr

📌 Backend Engineer — Model Inference (Ontario)
🏢 Best AI Tools Wiki
📍 Ontario

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: backend engineer — model inference (ontario) / ontario

Subscribe to this job alert:

Get the latest job offers by email for: backend engineer — model inference (ontario) / ontario