26 Aug
|
Veeda AI
|
Toronto
Join Veeda AI as a Member of Technical Staff specializing in high-performance computing. Drive cutting-edge innovations in AI infrastructure, focusing on GPU operations and advanced job scheduling. In this critical role, you'll work with a agile team tackling challenges at the frontier of Physical AI.
You'll be responsible for deploying and optimizing bare-metal GPU clusters, ensuring system reliability and efficiency through advanced troubleshooting and automation practices. Make a lasting impact on AI research and development from day one. Key Responsibilities:
- Deploy and operate GPU clusters using Kubernetes
- Fine-tune Slurm for optimal job performance
- Manage InfiniBand infrastructure and routing
- Ensure high-throughput data pathways for performance
- Build observability systems for cluster health monitoring Requirements:
- Bachelor’s degree in Computer Science or related field
- Extensive experience with Linux HPC systems
- Advanced diagnostics and troubleshooting capabilities
- Skilled in scripting and automation frameworks
- Hands-on experience with distributed AI workloads Your skills in HPC and AI will play a vital role at Veeda AI.
📌 Member of Technical Staff - HPC Specialist (Toronto)
🏢 Veeda AI
📍 Toronto