Contribute to groundbreaking AI solutions at Veeda AI as a Member of Technical Staff in infrastructure management. This role focuses on GPU clusters, efficiency, and system performance in a collaborative workplace. As a highly skilled engineer, you will be pivotal in managing and optimizing GPU infrastructure for advanced AI applications.
Your responsibilities will include automating deployment processes and implementing telemetry systems to enhance infrastructure reliability. This position is an excellent prospect for those eager to impact cutting-edge technology. Key Responsibilities:
Manage GPU cluster operations and configurations
Optimize job scheduling within Slurm and Kubernetes
Validate fabric engineering and performance tuning
Run high-throughput storage systems optimally
Construct telemetry pipelines for hardware observability Requirements:
Bachelor’s degree in a relevant field
Proficient in administering Linux HPC clusters
Expertise in troubleshooting hardware and network issues
Experience with infrastructure as code methodologies
Competency in tools for distributed deep learning Bring your advanced technical skills and passion for AI to Veeda AI’s cutting-edge team.
📌 Physical Ai Infrastructure Engineer Role Toronto
🏢 Veeda AI
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.