Talent.com
Advanced Micro Devices, Inc
Principal Software Developer – AI/ML Performance Validation & Systems TestingAdvanced Micro Devices, Inc • MARKHAM, Ontario, Canada
Principal Software Developer – AI/ML Performance Validation & Systems Testing

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, Inc • MARKHAM, Ontario, Canada
6 days ago
Job type
  • Full-time
Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.Together, we advance your career.

THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Create a job alert for this search

Principal Software Developer – AI/ML Performance Validation & Systems Testing • MARKHAM, Ontario, Canada

Similar jobs

Intact Senior AI Systems Developer

Intacttoronto, on, Canada
Full-time

As a Senior AI Systems Developer at Intact, you will shape the future of AI solutions.Join a flexible hybrid work setting and collaborate with domain experts.Intact is expanding and needs a Senior ... Show more

 • Promoted

Hybrid Principal Enterprise Architect - AI Focus

Priceline.com LLCToronto, Ontario, Canada
Full-time

Take on a leadership role as a Principal Enterprise Architect focusing on AI Platforms with a flexible hybrid work model.Drive the technical vision for AI and Machine Learning systems.This strategi... Show more

 • Promoted

Principal Staff Software Developer – AI/ML Performance Validation & Systems Testing

AMDMarkham, York Region, CA
Full-time

What you do at AMD changes everything.At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded syst... Show more

 • Promoted

Senior AI/ML Applications Architect

CB1173 GEPR Energy Canada IncMarkham, York region, Canada
Full-time

GE Vernova is accelerating the path to more reliable, affordable, and sustainable energy, while helping our customers power economies and deliver the electricity that is vital to health, safety, se... Show more

 • Promoted

AI Engineer

Frontier Dental CAMarkham, ON, Canada
Full-time

We’re seeking an<br/><br/>AI Engineer<br/><br/>to help shape the future of Frontier Dental through practical AI adoption and business automation.The key responsibilities and... Show more

 • Promoted

ML Platform Engineer at Mistplay

ODAIAToronto, Ontario, Canada
Full-time

Become a Principal Platform Engineer at Mistplay, focusing on ML infrastructure solutions while working hybrid in Toronto or Montreal.Create impactful real-time systems for model serving.In this ro... Show more

 • Promoted

Senior AI/ML Applications Architect

GE VernovaMarkham, Ontario, Canada
Full-time

GE Vernova is accelerating the path to more reliable, affordable, and sustainable energy, while helping our customers power economies and deliver the electricity that is vital to health, safety, se... Show more

 • Promoted

AI Developer Tools Engineer at Thumbtack

Thumbtacktoronto, on, Canada
Full-time

Elevate developer productivity at Thumbtack as a Senior Software Engineer specializing in AI tools.Contribute to our innovative engineering platform.As part of the Developer Experience team, you wi... Show more

 • Promoted

Principal Engineer (AI)

Resolver, a Kroll BusinessToronto, Ontario, Canada
Full-time

The Principal Engineer (AI) is the highest-ranking individual contributor, with deep expertise in artificial intelligence, including large language models, machine learning, generative AI, Agentic ... Show more

 • Promoted

Principal Ai Agent / Ml Software Engineer - C$81,700 - C$131,700 A Year

OracleToronto, Canada
Full-time

Lead AI/ML engineering for next-gen AI systems on OCI, focusing on scalable agentic AI platforms, autonomous workflows, and inference infrastructure. Show more

 • Promoted

Principal Software Engineer - AI Model Serving

Latinx in AI (LXAI)Toronto, Ontario, Canada
Full-time

Elevate your career at Workday as a Principal Software Development Engineer on the AI Model Serving team.Drive critical design decisions and lead high-impact projects in a flexible work environment... Show more

 • Promoted

Full Stack Developer with AI/ML Exposure

Trident StaffToronto
Full-time

Elevate your career as a Full Stack Developer in Toronto, specializing in Java, Spring Boot, and API development.This role is 100% onsite and ideal for experienced professionals.We are seeking a Fu... Show more

 • Promoted

AI Software Lead at Thomson Reuters

PowerToFlyToronto, Ontario, Canada
Full-time

Lead the development of cutting-edge AI software at Thomson Reuters.As a Lead Software Engineer, you'll play a crucial role in building automated solutions to revolutionize accounting practices in ... Show more

 • Promoted

AI/ML Solutions Lead at Iris Software

Iris Software Inc.toronto, on, Canada
Full-time

Accelerate your career as an AI/ML Solutions Lead in a hybrid position at Iris Software.Leverage your Python skills to drive innovative AI applications.As a lead engineer, you will be at the forefr... Show more

 • Promoted

Principal Engineer for AI Model Serving

HR Tech Jobtoronto, on, Canada
Full-time

Contribute to Workday's innovation as a Principal Engineer on the AI Model Serving team.Specialize in designing scalable distributed systems and guiding technical direction.This role involves shapi... Show more

 • Promoted

Senior AI Developer: Agentic Systems & Platform Lead

Constellation Dealer GroupToronto, ON, CA
Full-time

A North American software provider is seeking a Senior Developer to design and build AI-native systems that impact business outcomes.The role emphasizes collaborative system design and engineering ... Show more

 • Promoted

Agentic AI Systems Developer - Remote

NTT DATA, Inc.Toronto, ON, CA
Remote
Full-time

We are currently seeking a Agentic AI Systems Developer - Remote to join our team in Toronto, Ontario (CA-ON), Canada (CA).You will design and build agentic AI systems for healthcare using the Neur... Show more

 • Promoted

Vehicle Motion Control AI/ML Platform Design Engineer

General MotorsMarkham, ON, CA
Full-time

This posting is not for an existing vacancy within the organization and is open to new applications.As part of the application process, Artificial Intelligence will be used in the hiring process fo... Show more

 • Promoted

Lead AI Solutions Developer - CAA SCO

CAA Club GroupMarkham
Full-time

Join CAA SCO Systems & Services Inc.Lead Developer in AI Solutions.This hands-on role requires expertise in machine learning and AI application design.In this pivotal role at CAA SCO Systems & Serv... Show more

 • Promoted

AI Developer - LLM Features & AI Systems

STAN AIToronto, Ontario, Canada
Full-time

We're looking for an AI Developer to build the AI backbone of our product — retrieval-augmented generation pipelines, multi-step agent workflows, embedding systems, and LLM integrations that proper... Show more