Elevate your technical skills with Veeda AI as a Member of Technical Staff focused on AI infrastructure. Contribute to GPU cluster operations and optimize performance in a cutting-edge Physical AI environment.
Veeda AI is looking for a talented infrastructure engineer to join its fast-moving team. The ideal candidate will manage high-performance computing environments using Kubernetes and Slurm, while also tuning low-level networking and storage systems. This role offers the chance to significantly impact Physical AI by enabling advanced research and development efforts.
Key Responsibilities:
• Design and manage GPU clusters with Slurm/Kubernetes
• Tune job scheduling and QoS policies effectively
• Oversee the interconnect and validate it with tests
• Maintain high-throughput storage solutions seamlessly
• Develop telemetry pipelines for observability and health
Requirements:
• Bachelor's in Computer Science or equivalent experience
• Expertise in managing Linux HPC clusters
• Robust troubleshooting skills with hardware and networking
• Proficiency in automation tools like Ansible or Terraform
• Experience with deep learning frameworks such as PyTorch
Utilize your expertise in AI infrastructure to make meaningful contributions at Veeda AI.
#J-18808-Ljbffr
📌 AI Infrastructure Engineer at Veeda AI (Ontario)
🏢 Veeda AI
📍 Ontario
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.