06 Oct
|
Google Canada
|
London
06 Oct
Google Canada
London
Google Cloud is looking for a Technical Solutions Development Manager, Compute to join their team in Waterloo, Ontario. This is a senior-level technical leadership role at the intersection of AI/ML infrastructure, deep engineering expertise, and customer advocacy — ideal for someone who thrives on solving the most complex problems that enterprise customers encounter on large-scale cloud platforms.
In this role, you'll own complex customer escalations end to end, combining sharp technical troubleshooting skills with business acumen. You'll work across hardware and software boundaries, collaborate with Product and Site Reliability Development teams, and help shape the tools and processes that keep Google Cloud's AI infrastructure running at the highest levels of reliability.
About the Role: Technical Solutions Development Manager, Compute
As a Solutions Development Manager within Google Cloud's AI Infrastructure team, you'll take direct ownership of difficult customer issues — from initial diagnosis through resolution. That means reproducing customer-reported problems, determining root causes, and building tooling that accelerates future investigations. You'll be embedded in a global team delivering 24x7 support for customers deploying AI and ML workloads, so your ability to act decisively under pressure really matters here.
Beyond hands-on troubleshooting, this role carries a leadership dimension. You'll develop team members into highly skilled experts capable of diagnosing a wide variety of hardware and software boundary issues within minutes. You'll also serve as a consultant and subject matter expert for internal stakeholders across Development, Sales, and customer organizations — bridging technical depth with an understanding of what customers actually need to succeed with AI infrastructure deployments.
Benefits and Salary
This position offers a base salary of $174,000 – $178,000 CAD per year, plus a 20% bonus target, equity, and a full benefits package.
Google's benefits programme is comprehensive — visit Google's careers site to explore the full details of what's on offer.
Job Details
Company: Google
Location: Waterloo, ON, Canada
Pay: $174,000 – $178,000 CAD/year + 20% bonus target + equity + benefits
Responsibilities
This role sits at the centre of Google Cloud's AI Infrastructure support function, where technical ownership and cross-functional collaboration are equally important. Day to day, you'll be diagnosing complex issues, driving tooling improvements, and working alongside engineering and product teams to make the platform better for customers globally.
- Manage customer problems through effective diagnosis, resolution, or implementation of new investigation tools to increase productivity on AI/ML infrastructure issues
- Troubleshoot and reproduce customer-reported issues to determine root causes and build tools for faster diagnosis of AI/ML workloads and underlying hardware architectures
- Act as a consultant and subject matter expert for internal stakeholders in Development, Sales, and customer organizations to resolve complex deployment and operational obstacles in AI infrastructure environments
- Collaborate with Product and Development teams to identify product improvement opportunities, and work with Site Reliability Development teams to drive high-quality production outcomes
- Handle customer escalations by combining business acumen with technical skills across the hardware and software boundary
- Develop team members into highly skilled troubleshooting experts capable of diagnosing a wide variety of issues within minutes
- Lead operational excellence within the team with a focus on reliable execution and driving business growth by advocating for customers' AI deployment challenges
Requirements / Skills
Google is looking for a candidate with deep technical experience across AI/ML infrastructure, robust debugging and systems knowledge, and the leadership presence to mentor a high-performing team. You should be equally comfortable digging into kernel drivers and firmware as you are presenting to enterprise stakeholders.
- Bachelor's degree in Science, Technology, Engineering, Mathematics, or equivalent practical experience
- 13 years of experience reading and debugging code in general-purpose languages (e.g., Java, C, C++, Python, Shell, Go, or JavaScript) and in virtualization and orchestration frameworks
- Troubleshooting and triage experience across the full stack — hardware faults, low-level software, networking, virtualization, kernel drivers, firmware, and performance
- Linux/Unix systems expertise and experience debugging issues across the hardware/software boundary on enterprise-grade server infrastructure
- Experience with large-scale distributed systems and familiarity with common design patterns or best practices (preferred)
- Hands-on experience with AI/ML computing hardware, including GPUs or other accelerators (preferred)
- Familiarity with ML frameworks such as TensorFlow or PyTorch, and understanding of the AI/ML training and inference life-cycle (preferred)
- Knowledge of containerization and orchestration technologies like Kubernetes or Slurm in on-prem or cloud environments (preferred)
Bachelor's degree in Science, Technology, Engineering, Mathematics, or equivalent practical experience. 13 years of experience in reading/debugging code written in a general purpose coding language and in virtualization and orchestration frameworks. Experience troubleshooting and advocating for customer needs, and triaging technical issues across the stack. Experience with Linux/Unix systems and debugging issues across the hardware/software boundary on enterprise-grade server infrastructure.
#J-18808-Ljbffr
📌 Technical Solutions Development Manager, Compute, Google Cloud (London)
🏢 Google Canada
📍 London