31 Jul
|
Bagel Labs
|
Vancouver
31 Jul
Bagel Labs
Vancouver
We are Bagel Labs - a distributed machine learning research lab working towards open-source superintelligence. If you have high agency - meaning your default assumption is that you can control the outcome of whatever situation you are in - we want to hear from you. Every requirement below is adaptable for a candidate with high enough agency and tolerance for ambiguity.
You will design and optimize a distributed diffusion model training and serving system. Your focus is on building scalable, fault-tolerant infrastructure that can serve open-source diffusion models across multiple nodes and regions, with efficient support for adaptation techniques.
Design and implement distributed diffusion model inference systems for image, video, and multimodal generation across multiple nodes and regions.
Architect high-availability clusters for diffusion model serving with automatic failover, load balancing, and dynamic batching for variable-resolution outputs.
Build monitoring and observability systems for distributed diffusion inference (denoising steps, memory usage, generation latency, CLIP score tracking).
Integrate with open-source diffusion frameworks (Diffusers, ComfyUI, Invoke AI) and optimize for production-scale serving.
Implement and optimize cutting-edge techniques:
rectified flow models, consistency distillation, and progressive distillation for few-step generation.
Build infrastructure for effective LoRA/LyCORIS adaptation serving with hot-swapping and memory-efficient merging.
Optimize VAE decoding pipelines and implement tiled/windowed generation for ultra-high-resolution outputs.
Document architectural decisions, review code, and publish technical deep-dives on blog.At least 5 years of experience with distributed systems and production ML serving.
Deep understanding of diffusion model architectures (U-Net, DiT, rectified flows, consistency models).
Proven record of optimizing generation latency (classifier-free guidance, DDIM/DPM solvers, distillation techniques).
Experience with attention optimization techniques (Flash Attention, xFormers, memory-effective attention).
Strong understanding of adaptation techniques (LoRA, LyCORIS, textual inversion, DreamBooth).
A deeply technical culture where Full remote flexibility within North American time zones.
Ownership of work that can set the direction for decentralized AI.
Paid travel opportunities to the top ML conferences around the world.
📌 Machine Learning Engineer Fully Remote Vancouver
🏢 Bagel Labs
📍 Vancouver