Elevate your technical skills with Veeda AI as a Member of Technical Staff focused on AI infrastructure. Contribute to GPU cluster operations and optimize performance in a cutting-edge Physical AI setting.Veeda AI is looking for a talented infrastructure engineer to join its fast-moving team. The ideal candidate will manage high-performance computing environments using Kubernetes and Slurm, while also tuning low-level networking and storage systems. This role offers the chance to significantly impact Physical AI by enabling advanced research and development efforts.Key Responsibilities:
- Design and manage GPU clusters with Slurm/Kubernetes
- Tune job scheduling and QoS policies effectively
- Oversee the interconnect and validate it with tests
- Maintain high-throughput storage solutions seamlessly
- Develop telemetry pipelines for observability and healthRequirements:
- Bachelor's in Computer Science or equivalent experience
- Expertise in managing Linux HPC clusters
- Strong troubleshooting skills with hardware and networking
- Proficiency in automation tools like Ansible or Terraform
- Experience with deep learning frameworks such as PyTorchUtilize your expertise in AI infrastructure to make meaningful contributions at Veeda AI.#J-18808-Ljbffr
📌 Ai Infrastructure Engineer At Veeda Ai (Toronto)
🏢 Veeda AI
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.