Join Veeda AI as a Member of Technical Staff specializing in high-performance computing. Drive cutting-edge innovations in AI infrastructure, focusing on GPU operations and advanced job scheduling.In this critical role, you'll work with a agile team tackling challenges at the frontier of Physical AI. You'll be responsible for deploying and optimizing bare-metal GPU clusters, ensuring system reliability and efficiency through advanced troubleshooting and automation practices. Make a lasting impact on AI research and development from day one.Key Responsibilities:Deploy and operate GPU clusters using KubernetesFine-tune Slurm for optimal job performanceManage InfiniBand infrastructure and routingEnsure high-throughput data pathways for performanceBuild observability systems for cluster health monitoringRequirements:Bachelor's degree in Computer Science or related fieldExtensive experience with Linux HPC systemsAdvanced diagnostics and troubleshooting capabilitiesSkilled in scripting and automation frameworksHands-on experience with distributed AI workloadsYour skills in HPC and AI will play a vital role at Veeda AI.#J-18808-Ljbffr
📌 Member Of Technical Staff - Hpc Specialist (Toronto)
🏢 Veeda AI
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.