Elevate your technical skills with Veeda AI as a Member of Technical Staff focused on AI infrastructure. Contribute to GPU cluster operations and optimize performance in a cutting-edge Physical AI environment. Veeda AI is looking for a talented infrastructure engineer to join its rapid-moving team.
The ideal candidate will manage high-performance computing environments using Kubernetes and Slurm, while also tuning low-level networking and storage systems. This role offers the chance to significantly impact Physical AI by enabling advanced research and development efforts. Key Responsibilities:
- Design and manage GPU clusters with Slurm/Kubernetes
- Tune job scheduling and QoS policies effectively
- Oversee the interconnect and validate it with tests
- Maintain high-throughput storage solutions seamlessly
- Develop telemetry pipelines for observability and health Requirements:
- Bachelor's in Computer Science or equivalent experience
- Expertise in managing Linux HPC clusters
- Strong troubleshooting skills with hardware and networking
- Proficiency in automation tools like Ansible or Terraform
- Experience with deep learning frameworks such as PyTorch Utilize your expertise in AI infrastructure to make meaningful contributions at Veeda AI.
📌 AI Infrastructure Engineer at Veeda AI (Toronto)
🏢 Veeda AI
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.