Talent.com
Amazon Development Centre Canada ULC
Sr. SDE, Edge AI ML Platform, Edge AI and ScienceAmazon Development Centre Canada ULC • Vancouver, British Columbia, CAN
Les candidatures ne sont plus acceptées
Sr. SDE, Edge AI ML Platform, Edge AI and Science

Sr. SDE, Edge AI ML Platform, Edge AI and Science

Amazon Development Centre Canada ULC • Vancouver, British Columbia, CAN
Il y a 26 jours
Type de contrat
  • Temps plein
Description de poste

Amazon Devices (Lab126) builds products and services that delight millions of customers globally. The Edge AI ML Platform and Infrastructure team is building the platform that enables Amazon teams to train, optimize, evaluate, and deploy generative AI models on devices and in the cloud.

Today, optimizing a large model for a new hardware target requires experts to connect model onboarding, distributed training, compression, evaluation, compilation, and deployment systems by hand. We are turning that work into a repeatable, self-service workflow. Our platform supports large language, vision, audio, multimodal, and mixture-of-experts models. It gives scientists and engineers the tools to move new optimization techniques from research code into reliable production workflows.

We are looking for a Senior Software Development Engineer to lead the architecture and delivery of core ML platform capabilities. You will solve problems across distributed training on multi-node GPU clusters, model onboarding, compression pipelines, evaluation, GPU performance, artifact management, CI/CD, observability, and operational reliability. You will work with applied scientists, ML engineers, GPU kernel engineers, compiler and runtime teams, hardware teams, and product teams to deliver systems for models with hundreds of billions of parameters.

This role combines hands-on software development with technical leadership. You will write and review code, define architecture, resolve ambiguous requirements, lead projects that span multiple engineers and teams, and raise the engineering bar for an evolving ML platform.

Key job responsibilities
- Lead the design and delivery of distributed ML platform services and libraries across model ingestion, optimization, training, evaluation, packaging, and deployment.
- Define stable APIs and architecture boundaries that allow scientists to add algorithms without coupling research code to training, infrastructure, or deployment implementations.
- Design distributed training capabilities across data, tensor, pipeline, and model parallelism for large language and multimodal models.
- Scale workflows on multi-node GPU clusters while improving training throughput, GPU utilization, memory efficiency, communication performance, failure recovery, and developer iteration time.
- Develop infrastructure that connects distributed training with distillation, quantization, pruning, and other model optimization techniques.
- Build evaluation and artifact workflows that measure model quality and system performance, then carry validated models through deployment on target hardware.
- Build automated validation, CI/CD, regression testing, observability, and release mechanisms for GPU-intensive ML workloads.
- Profile and optimize end-to-end system performance with applied scientists and GPU kernel engineers. Translate bottlenecks into durable platform improvements.
- Establish operational mechanisms, including metrics, alarms, runbooks, on-call practices, and root-cause correction for production platform services.
- Partner with model, compiler, runtime, hardware, security, and infrastructure teams to clarify requirements, manage technical dependencies, and deliver multi-team programs.
- Write technical designs, evaluate trade-offs, and build consensus when the customer need is clear but the technology strategy is not.
- Mentor engineers, improve code and design review practices, and help recruit and develop a strong engineering team in Vancouver.

A day in the life
You will move between architecture and implementation. Your work will include reviewing designs for model onboarding interfaces, investigating failures in distributed training runs, profiling GPU workloads with scientists, leading cross-team reviews of end-to-end deployment paths, simplifying platform abstractions, and improving the release and regression mechanisms used by multiple model teams.

You will use performance, reliability, and developer productivity data to prioritize platform investments. You will make incremental deliveries while protecting long-term architecture, and you will ensure that the team resolves recurring problems at their root.

About the team
The Edge AI ML Platform and Infrastructure team brings together software engineers, ML infrastructure engineers, and GPU performance specialists. We build reusable model training, optimization, and deployment capabilities for Amazon product teams, working closely with applied scientists across Edge AI. Our customers need to adapt rapidly changing model architectures to constrained hardware and production workloads without rebuilding the toolchain for every model.

The team owns the platform foundations that connect model development to deployment. Our end-to-end scope lets us improve training, compression, evaluation, and deployment as one system. We value clear interfaces, measurable performance, automated quality gates, and direct collaboration between science and engineering.

BASIC QUALIFICATIONS

- 5+ years of non-internship professional software development experience
- 5+ years of programming with at least one software programming language experience
- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience as a mentor, tech lead or leading an engineering team
- Bachelor's degree in Computer Science, Engineering, or a related technical field
- Experience designing or building distributed systems or high-performance computing systems.

PREFERRED QUALIFICATIONS

- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Experience with CUDA kernels or ML/low-level kernels, or experience in debugging, profiling, and implementing software engineering best practices in large-scale systems
- Experience programming with at least one modern language such as Java, C++, or C# including object-oriented design, or experience with CUDA kernels or ML/low-level kernels
- Experience building distributed ML training, inference, evaluation, or data platforms using frameworks such as PyTorch, TensorFlow, JAX, NeMo, or Megatron.
- Experience with containers, Kubernetes, AWS infrastructure, CI/CD, observability, and production operations.
- Experience with model compression, quantization, knowledge distillation, model compilation, or edge deployment.
- Experience designing extensible platform APIs and delivering systems with science, hardware, compiler, or product teams.

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers.

Créer une alerte emploi pour cette recherche

Sr. SDE, Edge AI ML Platform, Edge AI and Science • Vancouver, British Columbia, CAN

Offres similaires

Remote AI Engineer - Build Next-Gen ML Solutions

Yeah! GlobalVancouver, Metro Vancouver Regional District, CA
Télétravail
Temps plein

A global technology firm is seeking highly skilled AI Engineers to join their team.This remote position requires proficiency in programming languages like Python and experience in developing AI mod... Voir plus

 • Offre sponsorisée

Remote SaaS Data & AI Engineering Lead

Harvey NashVancouver, Metro Vancouver Regional District, CA
Télétravail
Temps plein

A leading technology consultancy is seeking a Data & AI Engineering Manager to lead teams in building scalable data and AI products in a SaaS environment.The role combines technical leadership and ... Voir plus

 • Offre sponsorisée

Sr. Fullstack Engineer, Data and AI Platform

Rivian Automotive, Inc.Vancouver, Metro Vancouver Regional District, Canada
Temps plein

The Data and AI Platform team at Rivian supports developing and deploying Generative AI systems.We operate highly scalable infrastructure for serving and orchestrating LLMs, Vector Stores, and agen... Voir plus

 • Offre sponsorisée

Senior AI Engineer, AI Inference Systems

BrixPacific, British Columbia, Canada
Temps plein

Senior AI Engineer, AI Inference Systems.Senior AI Engineer, AI Inference Systems.Oversee and optimize the LLM inference engine to ensure performance, scalability, and cost-efficiency.Collaborate w... Voir plus

 • Offre sponsorisée

AI Operations Lead — Remote-First AI Transformation

commonskuVancouver, Metro Vancouver Regional District, CA
Télétravail
Temps plein

A tech-driven company is seeking an AI Operations Lead to spearhead internal AI transformation.This remote-first role focuses on enhancing workflows with AI while fostering a community-oriented env... Voir plus

 • Offre sponsorisée

Remote Forward-Deployed Engineer: AI-Driven Modernization

Banyan SoftwareVancouver, Metro Vancouver Regional District, CA
Télétravail
Temps plein

A leading software company is seeking a Forward Deployed Engineer to drive the modernization of applications with a focus on AI-assisted software development practices.This hybrid role requires ove... Voir plus

 • Offre sponsorisée

AI/ML Lead (Senior Machine Learning Engineer – Full-time Leadership Role)

Aurelian Venture AIvancouver, metro vancouver regional district, Canada
Temps plein

AI/ML Lead (Senior Machine Learning Engineer – Full-time Leadership Role).Full-Time, Contract, Hands‑on Technical Leadership.CAD $180,000 – $250,000 base + significant equity + performance bonuses ... Voir plus

 • Offre sponsorisée

AI Engineer

BioRenderVancouver, Metro Vancouver Regional District, CA
Temps plein

At BioRender, we don’t settle for average, and neither should you.We’re building tools that transform how scientists communicate complex ideas, and we need someone who’s ready to turn ambitious ide... Voir plus

 • Offre sponsorisée

Strategic Enterprise AE (West) - AI & Cloud Growth

UpboundVancouver, Metro Vancouver Regional District, Canada
Temps plein

A leading cloud technology company is looking for an Account Executive - Strategic Enterprise to manage and close business with major clients.The role involves collaborating with teams to enhance c... Voir plus

 • Offre sponsorisée

Agentic AI & Integration Technology Lead

Village Farms International Inc.Delta, Metro Vancouver Regional District, Canada
Temps plein

Take the lead in technology integration at Village Farms International, focusing on Agentic AI and enterprise workflows.This position offers a chance to ensure robust and secure integrations within... Voir plus

 • Offre sponsorisée

Lead AI and Full-Stack Engineer at Aequilibrium

AequilibriumVancouver, Metro Vancouver Regional District, Canada
Permanent

Step into a lead role at Aequilibrium focusing on AI-powered and full-stack engineering for immersive training solutions.Utilize your expertise to impact product outcomes and architectural integrit... Voir plus

 • Offre sponsorisée

Sr. UI Engineer

Omnissavancouver, metro vancouver regional district, Canada
Temps plein +1

The world is evolving fast, and organizations everywhere—from corporations to schools—are under immense pressure to provide flexible, work-from-anywhere solutions.They need IT infrastructure that e... Voir plus

 • Offre sponsorisée

Shape the Future of AI/ML with High-Impact Research Queries

Great Value HiringVancouver, Metro Vancouver Regional District, Canada
Temps plein

Engage as a pivotal AI/ML Research Expert by identifying game-changing questions in your field.Your work could lead to groundbreaking breakthroughs we need to tackle today.This role is crafted for ... Voir plus

 • Offre sponsorisée

Lead AI Platform Product Manager Role

Electronic Arts (EA)Vancouver, British Columbia, Canada
Temps plein

Drive AI product innovation at Electronic Arts as a Principal Product Manager.Your work will shape AI's role across game development and enhance user interaction.As part of the Infrastructure and P... Voir plus

 • Offre sponsorisée

Remote Engineering Manager, Data & AI Platform

ResonaiteVancouver, Metro Vancouver Regional District, CA
Télétravail
Temps plein

A fintech company is seeking an Engineering Manager to lead its Data & AI Engineering team.This role involves managing and developing engineering talent, establishing technical goals, and guiding d... Voir plus

 • Offre sponsorisée

Senior AI Developer: Agentic Systems & Platform Lead

Constellation Dealer GroupVancouver, Metro Vancouver Regional District, Canada
Temps plein

A North American software provider is seeking a Senior Developer to design and build AI-native systems that impact business outcomes.The role emphasizes collaborative system design and engineering ... Voir plus

 • Offre sponsorisée

Sanctuary AI Machine Learning Project Leadership

Sanctuary AIVancouver, Metro Vancouver Regional District, Canada
Temps plein

Join Sanctuary AI as a Technical Project Manager for groundbreaking machine learning initiatives.Drive project execution while facilitating Agile processes and ensuring high-quality deliverables.In... Voir plus

 • Offre sponsorisée

Senior AI Agent Engineer: Deploy & Integrate

CrestaVancouver, Metro Vancouver Regional District, CA
Temps plein

A leading AI technology company in Canada is seeking an AI Agent Engineer to develop and deploy intelligent AI agents.The ideal candidate will have a strong background in software development and A... Voir plus

 • Offre sponsorisée

AI Enablement Specialist

Rocky MountaineerVancouver
Temps plein

Professional & Business Support.Reporting to the Director, Corporate IT Services, the AI Enablement Specialist is the bridge between powerful AI tools and the people across our organization.The rol... Voir plus

 • Offre sponsorisée

Sr. Algorithm Engineer, AI

Comm100Vancouver
Temps plein

Research key capabilities leading to AGI, track the latest academic and industry research achievements in LLM, and bring new technical ideas and methods to the business.Leverage generative AI Agent... Voir plus