28 Aug
|
Adaption
|
Quebec City
28 Aug
Adaption
Quebec City
Adaption is seeking an engineer to own the cost and performance of our inference stack. Your work will determine how efficiently we serve models as workloads, traffic, and hardware change across our fleet.
You'll closely collaborate with the engineers operating the serving fleet, owning core levers such as caching, batching, quantization, decoding, and kernel-level optimization to boost throughput and reduce tail latency without compromising model quality.
#J-18808-Ljbffr
📌 Latency-Driven AI Inference Architect (Quebec City)
🏢 Adaption
📍 Quebec City