Senior ML Infra Engineer: Scale GPU Training & Reliability (Ontario)

Senior ML Infra Engineer: Scale GPU Training & Reliability (Ontario)

17 Aug
|
Veeda AI
|
Ontario

17 Aug

Veeda AI

Ontario

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We seek engineers to design and optimize distributed training systems across large GPU clusters, handling FP16/BF16/FP8 precision and debugging complex stability issues.
You will implement fault-detection, performance profiling, and resilient checkpointing to keep researchers productive in a fast-moving setting. Collaboration across teams is essential for success.

#J-18808-Ljbffr

📌 Senior ML Infra Engineer: Scale GPU Training & Reliability (Ontario)
🏢 Veeda AI
📍 Ontario

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior ml infra engineer: scale gpu training & reliability (ontario) / ontario

Subscribe to this job alert:

Get the latest job offers by email for: senior ml infra engineer: scale gpu training & reliability (ontario) / ontario