04 Aug
|
DRW Holdings
|
Montreal
04 Aug
DRW Holdings
Montreal
Join DRW as a DevOps HPC Specialist and work on optimizing GPU infrastructures for cutting-edge AI and machine learning systems. Be part of a cooperative team that thrives on innovation. As an HPC Specialist at DRW, based in Chicago, you will focus on deploying and maintaining GPU infrastructures designed to support advanced AI and ML workloads.
You’ll use your extensive experience in infrastructure engineering to optimize performance and scalability, ensuring efficient operation across complex systems. Collaboration with machine learning engineers will be key to improving model performance. Key Responsibilities:
- Optimize GPU infrastructure for LLM inference
- Manage multi-GPU deployments in Kubernetes
- Configure and manage network components
- Implement storage solutions for model efficiency
- Troubleshoot technical performance issues Requirements:
- Bachelor's or Master's in Computer Science
- Over 5 years of experience in infrastructure roles
- Strong skills in GPU management and optimization
- Proficient with Ansible or similar tools
- Excellent understanding of distributed systems Leverage your expertise in high-performance computing to deliver cutting-edge solutions at DRW.
📌 DevOps HPC Specialist at DRW (Montreal)
🏢 DRW Holdings
📍 Montreal