Talent.com
Amazon Devices
ML Kernel Performance Engineer, Edge AI and ScienceAmazon Devices • Vancouver, British Columbia, Canada
ML Kernel Performance Engineer, Edge AI and Science

ML Kernel Performance Engineer, Edge AI and Science

Amazon Devices • Vancouver, British Columbia, Canada
30+ days ago
Salary
CA$114,800.00 yearly
Job type
  • Full-time
Job description
Amazon Devices is an inventive research and development company that designs and engineers high-profile consumer products like the Kindle family Fire Tablets Fire TV Health & Wellness devices Amazon Echo and Astro. We are building the next generation of edge AI capabilities through our advanced compression platform and custom neural accelerator silicon.

Within Edge AI & Science the AI Platform team builds a compression platformthe first of its kindenabling 20-100x neural network compression for edge and cloud deployment. As model sizes grow from billions to hundreds of billions of parameters compute efficiency becomes the single largest return on engineering investment during training. The gap between eager-mode Python and optimized GPU execution is where months of training time are won or lost.

We are looking for an ML Kernel Performance Engineer to work at the hardware-software boundary of this platform crafting high-performance CUDA and Triton kernels that make our compression algorithms run at peak efficiency during training fine-tuning and inference. You will build the tooling and kernel libraries that democratize GPU performance optimization across the team enabling scientists and engineers to profile diagnose and fix kernel bottlenecks without needing to be CUDA experts themselves.

Working alongside compression scientists and platform engineers you will ensure that novel quantization schemes (ternary nonary mixed-precision) and sparse computation patterns translate into real throughput gains on GPU hardware. Your work will directly accelerate every training run in the organization and unlock deployment of compressed models to both edge devices and cloud inference.

Key job responsibilities
Design and implement high-performance CUDA and Triton kernels for quantization-aware training sparse matrix operations and low-bit inference on modern GPU accelerators

Analyze and optimize kernel-level performance for compression training workloads conducting detailed performance analysis using profiling tools to identify and resolve bottlenecks that slow model training from days to weeks

Implement kernel-level optimizations such as operator fusion tiling memory access pattern optimization and scheduling for compression-specific compute patterns

Build a kernel development harness that enables any team member to profile kernel performance test forward/backward accuracy and validate at production scale lowering the bar from CUDA expert to any engineer with agents

Maintain and extend the teams training kernels library with clean interfaces CI and examples that enable scientists to contribute kernel improvements alongside platform engineers

Collaborate closely with Applied Scientists compiler engineers and hardware architects to co-design ML-centric solutions that unify software and hardware for both cloud and edge deployment

Develop inference kernels for cloud deployment (custom backends for quantized models that keep weights packed in memory and reconstruct on the fly for compute)

Build and maintain performance regression tests and benchmarking infrastructure that track kernel efficiency as models scale from billions to hundreds of billions of parameters

A day in the life
A scientist files a ticket: QAT training on our large model is 4x slower than expected. You pull up the profiler identify that a custom quantizer kernel is thrashing shared memory at scale write a Triton replacement that tiles correctly for the layer shapes at that model size validate accuracy in the test harness and push it to the kernels repo. By end of day the training run that was taking four days now takes one.

You will also build the tooling that makes this workflow repeatable by others. You will participate in design discussions with Applied Scientists translate their algorithmic ideas into efficient GPU implementations and work in a startup-like environment where every engineering hour directly accelerates the teams ability to ship compressed models.

About the team
The AI Platform team builds Amazons neural network compression platform. We compress models using knowledge distillation network restructuring and advanced quantization to achieve 20-100x compression while preserving model quality. Our platform packages these into automated pipelines that deploy to both custom edge silicon and GPU-based cloud inference.

As model sizes grow the proprietary advantage shifts from the science to the software (making it work at hundreds of billions of parameters is the moat). GPU kernel performance is the biggest single lever on training throughput and we expect AI-assisted development tooling to significantly multiply engineering productivity meaning a small team with the right harness can operate at the scale of a much larger one.

The ML Kernel Performance Engineer bridges science and platforms: you turn algorithmic innovations into production-grade GPU code that runs at scale. You will work alongside Applied Scientists compiler engineers hardware architects and platform developers in a small agile team building the next generation of edge AI for Amazons consumer products.

- 3 years of non-internship professional software development experience
- 2 years of non-internship design or architecture (design patterns reliability and scaling) of new and existing systems experience
- Experience with CUDA kernels or ML/low-level kernels or experience in developing and deploying LLMs in production on GPUs Neuron TPU or other AI acceleration hardware
- Experience with programming languages such as Python Java C

- Bachelors degree in computer science or equivalent
- 3 years of full software development life cycle including coding standards code reviews source control management build processes testing and operations experience
- Experience with GPU kernel optimization and GPGPU computing (CUDA Triton SYCL or ROCm)
- Proficiency in low-level performance optimization for GPUs
- Understanding of GPU memory hierarchies and optimization strategies (shared memory L1/L2 cache register pressure memory coalescing)
- Experience developing high-performance libraries for ML or HPC applications
- Knowledge of ML frameworks (PyTorch TensorFlow) and their GPU backends
- Experience implementing custom PyTorch operators ( C extensions)
- Experience with parallel programming and optimization techniques
- Background in neural network compression (quantization pruning knowledge distillation low-rank factorization)
- Knowledge of mixed-precision training and inference (FP16 BF16 FP8 INT8 INT4)
- Experience with inference optimization (TensorRT ONNX Runtime vLLM or similar)
- Familiarity with Transformer architectures attention mechanisms and their compute/memory profiles
- Experience with AWS Trainium/Inferentia or the Neuron Kernel Interface (NKI)
- Experience with edge deployment model compilation or hardware-aware optimization

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status disability or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process including support for the interview or onboarding process please visit for more information. If the country/region youre applying in isnt listed please contact your Recruiting Partner.

The base salary range for this position is listed below. As a total compensation company Amazons package may include other elements such as sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience qualifications and location. Amazon offers comprehensive benefits including health insurance (medical dental vision prescription basic life & AD&D insurance) Registered Retirement Savings Plan (RRSP) Deferred Profit Sharing Plan (DPSP) paid time off and other resources to improve health and well-being. We thank all applicants for their interest however only those interviewed will be advised as to hiring status.



CAN BC Vancouver - 114800.00 - 191800.00 CAD annually


Required Experience:

IC


Employment Type : Full-Time
Department / Functional Area: Software Development
Experience: years
Vacancy: 1
Create a job alert for this search

ML Kernel Performance Engineer, Edge AI and Science • Vancouver, British Columbia, Canada

Similar jobs

Machine Learning Engineer, AICE - AI Center of Excellence

AmazonVancouver, British Columbia, Canada
Full-time

The AI Center of Excellence (AICE) builds AI primitives that power system‑intrinsic intelligence and Trusted Intelligent Knowledge Infrastructure for both AI‑as‑a‑Consumer and human users.We develo... Show more

 • Promoted

Geospatial Ml Engineer: Build & Deploy Ai Pipelines - C$110,000 - C$130,000 A Year

Scientific Testing FirmVancouver, Canada
Full-time

Seeking a Machine Learning Engineer to design and maintain ML systems, develop microservices, optimize performance, and collaborate with teams.Requires Python, AWS, ML, and data engineering skills. Show more

 • Promoted

ML Engineer (Speech-to-Speech) — Subject Matter Expert

VosynVancouver, Metro Vancouver Regional District, Canada
Full-time

ML Engineer (Speech-to-Speech) — Subject Matter Expert.At Vosyn, we embrace the exciting, game-changing world of Artificial Intelligence, driving innovation and pioneering impactful projects across... Show more

 • Promoted

Senior ML & CV Scientist - Edge AI for Robotics (Vancouver)

Apera AI Incvancouver, metro vancouver regional district, Canada
Full-time

A leading AI robotics company in Vancouver seeks a Senior Machine Learning / Computer Vision Applied Scientist.The role involves developing state-of-the-art robotic vision systems using deep learni... Show more

 • Promoted

Ml Vision Engineer — Edge & Cloud Ai - C$84,300 - C$120,000 A Year

Leading Mining Technology FirmVancouver, Canada
Full-time

Seeking a Machine Learning Developer to design AI algorithms for embedded, edge, and cloud environments.Requires Master's degree and 3+ years of experience. Show more

 • Promoted

Machine Learning Engineer/Senior Machine Learning Engineer - Evisort

HR Tech Jobvancouver, metro vancouver regional district, Canada
Full-time

As a Machine Learning Engineer, you will help develop tailored user experiences using advanced LLMs, Knowledge Graphs, personalization, and predictive analysis.You will collaborate with other engin... Show more

 • Promoted

Staff ML Engineer: Edge AI for Global Ops (Remote)

SamsaraVancouver, Metro Vancouver Regional District, Canada
Remote
Full-time

A technology company is looking for a Staff Machine Learning Engineer to work on transformative AI solutions.Based in Canada, this remote role involves leading AI product initiatives and collaborat... Show more

 • Promoted

Senior Developer (AI/ML/Gen AI Solutions)

Intello Technologies Inc.Burnaby
Full-time

Select how often (in days) to receive an alert:.Senior Developer (AI/ML/Gen AI Solutions).Location: Burnaby, BC, CA Ottawa, ON, CA Calgary, AB, CA Toronto, ON, CA Edmonton, AB, CA Vancouver, BC, CA... Show more

 • Promoted

Machine Learning Engineer

ClioVancouver, British Columbia, Canada
Full-time

Clio is the global leader in legal AI technology, empowering legal professionals and law firms of every size to work smarter, faster, and more securely.We are transforming the legal experience for ... Show more

 • Promoted

Kernel Engineer

Acceler8 TalentVancouver, Metro Vancouver Regional District, CA
Full-time

Kernel Engineer (AI Accelerator).We are seeking a full remote Kernel Engineer to join a heavily funded ($150m+ Series B) team building next generation hardware AI accelerators for neural net infere... Show more

 • Promoted

Senior Backend Engineer, AI Platform & MLOps

CrestaVancouver, Metro Vancouver Regional District, CA
Full-time

A leading AI solutions company is hiring an ML Engineer in Toronto, Canada.In this role, you will design and maintain serving stacks for machine learning models and automate training pipelines usin... Show more

 • Promoted

Senior Machine Learning Engineer

InstacartVancouver, Metro Vancouver Regional District, CA
Permanent

We're transforming the grocery industry.At Instacart, we invite the world to share love through food because we believe everyone should have access to the food they love and more time to enjoy it t... Show more

 • Promoted

AI ML Architect

3PillarVancouver, Metro Vancouver Regional District, CA
Full-time

At 3Pillar, culture is more than a buzzword.The power of culture, teamwork, and open collaboration drive our commitment to building breakthrough software solutions that power digital businesses.Our... Show more

 • Promoted

Ml Vision Engineer — Edge & Cloud Ai - C$84,300 - C$120,000 A Year

A leading mining technology firmVancouver, Canada
Full-time

Machine Learning Developer sought for innovative AI projects, designing algorithms for embedded, edge, and cloud environments. Show more

 • Promoted

Geospatial Ml Engineer: Build & Deploy Ai Pipelines - $110,000 - $130,000 A Year

AlsglobalVancouver, Canada
Full-time

Machine Learning Engineer needed to design and maintain ML systems, develop microservices, and optimize performance. Show more

 • Promoted

Machine Learning Engineer - Diffusion

Bagel LabsVancouver, Metro Vancouver Regional District, Canada
Full-time

We are Bagel Labs - a distributed machine learning research lab working towards open-source superintelligence.We ignore years of experience and pedigree.If you have high agency - meaning your defau... Show more

 • Promoted

Staff Machine Learning Engineer - Generative Ai Platform - $135,000 - $204,300 A Year

InfobloxBurnaby, Canada
Full-time

Designs and deploys enterprise-grade Generative AI systems, including product copilots and agentic AI solutions, to power the next generation of networking and security experiences. Show more

 • Promoted

Senior ML Inference Engineer — Model Efficiency

CohereVancouver, Metro Vancouver Regional District, CA
Full-time

A leading AI technology company is seeking a Member of Technical Staff to enhance model efficiency.This role involves improving performance metrics, optimizing bottlenecks, and collaborating with v... Show more

 • Promoted

Senior Ml Engineer — Production Ai Agents - C$188,000 - C$282,000 A Year

AI platform companyVancouver, Canada
Full-time

Designs and builds core ML systems for AI agents, requiring expertise in applied ML, PyTorch, TensorFlow, and cloud computing. Show more

 • Promoted

Machine Learning Engineer

BioRenderVancouver, Metro Vancouver Regional District, CA
Full-time

At BioRender, we are accelerating the world’s ability to discover, learn, and communicate science faster through visuals.Today, BioRender empowers millions of scientists to create beautiful, accura... Show more