AI/ML Solutions Architect – Distributed Training & GPU InfrastructureLocation: Remote from anywhere CanadaTotal compensation up to ~$400k (base + variable), depending on level and experienceJoin a fast-moving AI infrastructure team working on the cutting edge of large-scale ML workloads. This role is ideal for engineers who enjoy solving deep technical challenges in distributed training, multi-GPU systems, and scalable AI inference infrastructure. You will work directly with AI-focused clients, helping them get the most out of contemporary GPUs (H100, B200, etc.) and ML frameworks such as PyTorch (and JAX in some environments).Team & ResponsibilitiesWork alongside senior AI and infrastructure engineers building large-scale GPU platforms. As part of the customer solutions team, you will:Design and validate production-grade distributed training (primary) and large-scale inference architectures on large GPU clusters, typically tens to thousands of GPUsWork hands-on with customers to debug, optimize, and scale ML workloads across multi-node GPU environmentsAct as a technical authority on GPU performance, networking, and schedulers,
making trade-offs at scale and translating customer needs into concrete platform requirementsCollaborate closely with engineering, product, and R&D to influence roadmap decisions based on real-world ML workloadsThis is a hands-on, technical role ; you are expected to work directly in customer environments, not only advise at a high levelRequired SkillsHands-on experience designing and operating production-grade, multi-node GPU workloads for training or inferenceStrong background in distributed deep learning (PyTorch Distributed, DeepSpeed) on GPU clustersDeep understanding of GPU architecture and interconnects (H100/A100 class, NVLink, InfiniBand)Experience with Kubernetes or Slurm and performance tuning using GPU profiling and monitoring toolsThis role is not a fit if your experience is limited to single-node training, high-level AI strategy, or non-production research environments. We are looking for engineers and architects who thrive at the intersection of AI workloads and large-scale infrastructure.#J-18808-Ljbffr
📌 Ai/Ml Solution Architect - Up To $300,000 A Year - Remote (Toronto)
🏢 Doghouse Recruitment
📍 Toronto
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.