25 Aug
|
Veeda AI
|
Toronto
Elevate your technical skills with Veeda AI as a Member of Technical Staff focused on AI infrastructure. Contribute to GPU cluster operations and optimize performance in a cutting-edge Physical AI setting. Veeda AI is looking for a talented infrastructure engineer to join its rapid-moving team.
The ideal candidate will manage high-performance computing settings using Kubernetes and Slurm, while also tuning low-level networking and storage systems. This role offers the chance to significantly impact Physical AI by enabling advanced research and development efforts. Key Responsibilities:
Design and manage GPU clusters with Slurm/Kubernetes
Tune job scheduling and QoS policies effectively
Oversee the interconnect and validate it with tests
Maintain high-throughput storage solutions seamlessly
Develop telemetry pipelines for observability and health Requirements:
Bachelor's in Computer Science or equivalent experience
Expertise in managing Linux HPC clusters
Strong troubleshooting skills with hardware and networking
Proficiency in automation tools like Ansible or Terraform
Experience with deep learning frameworks such as PyTorch Utilize your expertise in AI infrastructure to make meaningful contributions at Veeda AI.
📌 Ai Infrastructure Engineer At Veeda Ai Toronto
🏢 Veeda AI
📍 Toronto