Latency-Driven AI Inference Architect (Winnipeg)

Latency-Driven AI Inference Architect (Winnipeg)

29 Aug
|
Adaption
|
Winnipeg

29 Aug

Adaption

Winnipeg

Adaption is seeking an engineer to own the cost and performance of our inference stack. Your work will determine how efficiently we serve models as workloads, traffic, and hardware change across our fleet. You'll closely collaborate with the engineers operating the serving fleet, owning core levers such as caching, batching, quantization, decoding, and kernel-level optimization to boost throughput and reduce tail latency without compromising model quality.

#J-18808-Ljbffr

📌 Latency-Driven AI Inference Architect (Winnipeg)
🏢 Adaption
📍 Winnipeg

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: latency-driven ai inference architect (winnipeg) / winnipeg