Latency-Driven AI Inference Architect (Quebec City)

Latency-Driven AI Inference Architect (Quebec City)

28 Aug
|
Adaption
|
Quebec City

28 Aug

Adaption

Quebec City

Adaption is seeking an engineer to own the cost and performance of our inference stack. Your work will determine how efficiently we serve models as workloads, traffic, and hardware change across our fleet.
You'll closely collaborate with the engineers operating the serving fleet, owning core levers such as caching, batching, quantization, decoding, and kernel-level optimization to boost throughput and reduce tail latency without compromising model quality.

#J-18808-Ljbffr

📌 Latency-Driven AI Inference Architect (Quebec City)
🏢 Adaption
📍 Quebec City

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: latency-driven ai inference architect (quebec city) / quebec city

Subscribe to this job alert:

Get the latest job offers by email for: latency-driven ai inference architect (quebec city) / quebec city