Principal Software Developer - AI/ML Performance Validation & Systems Testing (Markham)

Principal Software Developer - AI/ML Performance Validation & Systems Testing (Markham)

19 Aug
|
AMD
|
Markham

19 Aug

AMD

Markham

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Join us as we shape the future of AI and beyond. We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems.

In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship— from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production. Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.

Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them. Architect the test infrastructure — distributed test runners, GitHub Actions/ Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines. Champion modern, agile quality engineering — shift-left testing, test pyramids,



contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation intrunk.

Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects. Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage. Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.

Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance. Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives. Lead system-level testing for server nodes — multi-GPU topologies, PCIe/InfinityFabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.

Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks— establishing reproducible methodology, baselines, and regression tracking. Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.

Expert





Python for test automation and infrastructure; strong C++ for debugging and extending production code. GPU software stacks (ROCm, CUDA, oneAPI, SYCL) AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM) Linux kernel, GPU drivers, or accelerator firmware Distributed systems and large-scale cluster software Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers. Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.

Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.

Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics. Familiarity with AI training, inference, and HPC benchmark methodologies.

Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis. Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.

ACADEMIC CREDENTIALS: ~ BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience). AMD and its subsidiaries are equal prospect, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

📌 Principal Software Developer - AI/ML Performance Validation & Systems Testing (Markham)
🏢 AMD
📍 Markham

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal software developer - ai/ml performance validation & systems testing (markham) / markham

Subscribe to this job alert:

Get the latest job offers by email for: principal software developer - ai/ml performance validation & systems testing (markham) / markham