Talent.com
Advanced Micro Devices, Inc
Principal Software Developer – AI/ML Performance Validation & Systems TestingAdvanced Micro Devices, Inc • MARKHAM, Ontario, Canada
Principal Software Developer – AI/ML Performance Validation & Systems Testing

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, Inc • MARKHAM, Ontario, Canada
Il y a 1 jour
Type de contrat
  • Temps plein
Description de poste

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.

THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Créer une alerte emploi pour cette recherche

Principal Software Developer – AI/ML Performance Validation & Systems Testing • MARKHAM, Ontario, Canada

Offres similaires

Ai Performance Architect: Hardware-Software Co-Design (Remote) - $100,000 - $300,000 A Year - Remote

TenstorrentToronto, Canada
Télétravail
Temps plein

AI Performance Architect sought to model and optimize AI workloads.Requires C++ and Python experience.Salary ranges from $100k to $500k. Voir plus

 • Offre sponsorisée

Senior Ml Engineer - Ai Platform & Ranking Systems - C$160,000 - C$220,000 A Year

J-18808-LjbffrToronto County, Canada
Temps plein

Seeking a Senior Machine Learning Engineer to develop and deploy ML solutions, focusing on recommendation systems, for a relationship intelligence platform. Voir plus

 • Offre sponsorisée

Hybrid Principal Enterprise Architect - AI Focus

Priceline.com LLCToronto, Ontario, Canada
Temps plein

Take on a leadership role as a Principal Enterprise Architect focusing on AI Platforms with a flexible hybrid work model.Drive the technical vision for AI and Machine Learning systems.This strategi... Voir plus

 • Offre sponsorisée

Senior Principal Researcher – AI Agent & Multimodal Interaction System

Huawei CanadaMarkham, ON, CA
Permanent

Huawei Canada has an immediate permanent opening for a Senior Principal Researcher.The Huawei Human-Machine Interaction Lab unites global researchers, engineers, and designers to redefine human tec... Voir plus

 • Offre sponsorisée

Senior Developer (Ai/Ml/Gen Ai Solutions) - C$104,000 - C$154,000 A Year

TelusToronto, Canada
Temps plein

Seeking a Senior Full Stack Developer with AI/ML/Gen AI expertise to lead cross-functional teams in designing and implementing intelligent applications, mentor junior developers, and drive technica... Voir plus

 • Offre sponsorisée

Senior Principal Researcher – AI Agent & Multimodal Interaction System

Huawei Technologies Canada Co., Ltd.Markham, ON, CA
Permanent

Huawei Canada has an immediate permanent opening for a Senior Principal Researcher.The Huawei Human-Machine Interaction Lab unites global researchers, engineers, and designers to redefine human tec... Voir plus

 • Offre sponsorisée

Principal Software Developer — AI-Driven, Scalable Solutions

Intuit Inc.Toronto, ON, CA
Temps plein

A leading software company is seeking a Principal Software Developer in Toronto to drive technology initiatives and collaborate on customer solutions.Candidates should have full-stack development e... Voir plus

 • Offre sponsorisée

Senior Applied AI Engineer — Product & Revenue Systems

QEA TechMarkham, ON, CA
Temps plein

Build and deploy AI systems that directly drive revenue by embedding AI across QEA’s product, sales, and marketing workflows.This role focuses on execution: turning QEA’s building data, imagery, an... Voir plus

 • Offre sponsorisée

Senior Applied Ai Systems Developer / Architect - C$120,000 - C$150,000 A Year

Klick GroupEast York, Canada
Temps plein

Senior developer/architect to lead AI-enabled features for the Genome platform, focusing on scalable systems and integrating AI/LLM capabilities. Voir plus

 • Offre sponsorisée

Senior Ai Developer: Agentic Systems & Platform Lead - C$150,000 - C$185,000 A Year

Constellation Dealer GroupNorth York, Canada
Temps plein

Lead the design and development of AI-native systems, focusing on agentic coding practices and efficient workflows.Requires strong backend and distributed systems skills. Voir plus

 • Offre sponsorisée

Senior Ml Engineer - Ai Platform & Ranking Systems - C$160,000 - C$220,000 A Year

Affinity.coToronto, Canada
Temps plein

A leading relationship intelligence platform in Toronto is seeking a Senior Machine Learning Engineer to advance their ML engineering capabilities.This role involves owning the ML lifecycle and tra... Voir plus

 • Offre sponsorisée

Principal Software Architect — AI-Driven Data Platforms Lead

AlphaSenseToronto, ON, CA
Temps plein

A leading market intelligence firm in Canada seeks an experienced engineer to own the architecture and evolution of a large-scale data extraction platform.The role involves designing reliable syste... Voir plus

 • Offre sponsorisée

Senior AI/ML Applications Architect

GE VernovaMarkham, York Region, CA
Temps plein

GE Vernova is accelerating the path to more reliable, affordable, and sustainable energy, while helping our customers power economies and deliver the electricity that is vital to health, safety, se... Voir plus

 • Offre sponsorisée

Principal Enterprise Architect, AI Platform

Priceline.com LLCToronto, Ontario, Canada
Temps plein

Principal Enterprise Architect, AI Platform.Location: Hybrid (Two days in-office) Overview.The Principal Enterprise Architect for the AI Platform is a high‑impact leadership role responsible for se... Voir plus

 • Offre sponsorisée

Senior AI Software Engineer - On-Device ML (C++17)

QualcommMarkham, York Region, CA
Temps plein

A leading technology company in York Region, Markham is seeking a software engineer to develop AI solutions using modern C++17.You will work on cutting-edge technology for Qualcomm Hexagon Processo... Voir plus

 • Offre sponsorisée

AI Platform Software Engineer

GeotabToronto, Ontario, Canada
Temps plein

Empower AI initiatives as a Senior Data Platform Developer focused on machine learning technologies.Develop, maintain, and enhance AI applications while ensuring the performance and reliability of ... Voir plus

 • Offre sponsorisée

Senior Ml Engineer - Ai Platform & Ranking Systems - C$160,000 - C$220,000 A Year

Relationship Intelligence PlatformNorth York, Canada
Temps plein

Seeking a Senior Machine Learning Engineer to lead ML engineering efforts, focusing on the ML lifecycle and building recommendation systems.Competitive salary and benefits offered. Voir plus

 • Offre sponsorisée

AI Developer Tools Engineer at Thumbtack

ThumbtackToronto, ON, CA
Temps plein

Elevate developer productivity at Thumbtack as a Senior Software Engineer specializing in AI tools.Contribute to our innovative engineering platform.As part of the Developer Experience team, you wi... Voir plus

 • Offre sponsorisée

Senior AI Developer: Agentic Systems & Platform Lead

Constellation Dealer GroupToronto, ON, CA
Temps plein

A North American software provider is seeking a Senior Developer to design and build AI-native systems that impact business outcomes.The role emphasizes collaborative system design and engineering ... Voir plus

 • Offre sponsorisée

Remote AI & ML Developer — LLMs, RAG & APIs

OrganimiToronto, ON, CA
Télétravail
Temps plein

A leading B2B SaaS company is seeking an experienced AI / Machine Learning Developer to work on innovative AI models and systems.The candidate will train and deploy large language models and relate... Voir plus