Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. ML Kernel Performance Engineer, AWS Neuron

$193.5k
Full-time

Annapurna Labs (U.S.) Inc.

Salary: $193,500 - 193,500 per year Requirements:

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming experience in at least one software language
  • 5+ years of experience leading design or architecture for new and existing systems, including design patterns, reliability, and scaling
  • 5+ years of full software development life cycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations
  • Experience serving as a mentor, tech lead, or engineering team lead
  • Bachelors degree in computer science or an equivalent background
  • 6+ years of full software development experience
  • Expertise in accelerator architectures for ML or HPC, such as GPUs, CPUs, FPGAs, or custom architectures
  • Experience with GPU kernel optimization and GPGPU computing, such as CUDA, NKI, Triton, OpenCL, SYCL, or ROCm
  • Demonstrated experience with NVIDIA PTX and/or AMD GPU ISA
  • Experience building high-performance libraries for HPC applications
  • Proficiency in low-level performance tuning for GPUs
  • Experience developing LLVM/MLIR backends for GPUs
  • Knowledge of ML frameworks such as PyTorch or TensorFlow and their GPU backends
  • Experience with parallel programming and optimization methods
  • Understanding of GPU memory hierarchies and optimization strategies
Responsibilities:
  • Design and build high-performance compute kernels for ML operations using the Neuron architecture and programming models
  • Analyze and improve kernel-level performance across multiple generations of Neuron hardware
  • Use profiling tools to perform detailed performance analysis and identify bottlenecks
  • Develop compiler optimizations including fusion, sharding, tiling, and scheduling
  • Partner directly with customers to enable and optimize their ML models on AWS accelerators
  • Work across teams to create innovative kernel optimization techniques
  • Architect and implement business-critical features
  • Publish cutting-edge research
  • Mentor experienced engineers
Technologies:
  • AWS
  • Architect
  • CUDA
  • Hardware
  • LLVM
  • Machine Learning
  • PyTorch
  • TensorFlow
  • Web
  • Cloud
  • AI
  • Backbone
  • Backend
  • Flow
  • Support

More:

We are the Annapurna Labs team at Amazon Web Services, building AWS Neuron, the software development kit that accelerates deep learning and GenAI workloads on our custom machine learning accelerators, Inferentia and Trainium. Our Acceleration Kernel Library team focuses on maximizing performance at the hardware-software boundary, crafting high-performance kernels for ML functions so every FLOP counts for customer workloads. We work across the Neuron Compiler organization, spanning frameworks, compilers, runtime, and collectives, while also contributing to future architecture designs and collaborating closely with customers on model enablement. This role offers the chance to work on cutting-edge products at the intersection of machine learning, high-performance computing, and distributed architectures. We value an innovative, agile, and collaborative culture, and we offer flexible working arrangements, a strong work-life balance, inclusive team programs, and comprehensive benefits including health coverage, retirement matching, paid time off, parental leave, and sign-on payments and RSUs. The position is based in Cupertino, California, with a listed annual base salary of 193,500.00 USD.

last updated 37 week of 2026

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Sr. ML Kernel Performance Engineer, AWS Neuron in Cupertino, CA vacancy
  • $193.3k - $261.5k

     ...at Amazon Web Services (AWS) builds AWS Neuron, the software development...  ...Trainium.The Acceleration Kernel Library team is at the forefront of maximizing performance for AWS's custom ML accelerators. Working at...  ...hardware-software boundary, our engineers craft high-performance... 
    Amazon Web Service
    Senior
    Performance
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $165.6k

     ...Responsibilities: Design and build high-performance compute kernels for machine learning operations using the Neuron architecture and programming...  ...machine learning models on AWS accelerators Partner with...  ...get the most out of AWS ML accelerators. We operate across... 
    Amazon Web Service
    Performance
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 day ago
  • $253.1k - $342.3k

    Utility Computing (UC)AWS Utility Computing (...  ...that builds AWS Neuron, the software stack...  ...accelerators.As a Sr. Software Development...  ...and with ML model developers and...  ...SDK to deliver best performance and usability on top...  ...qualifications- 10+ years of engineering experience- 5+... 
    Amazon Web Service
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom...  ...fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team...  ...job responsibilities* Deliver high-performance models using distributed inference... 
    Amazon Web Service
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $193.3k - $261.5k

     ...as a tech lead, or leading an engineering team ~ Masters degree in...  ...Responsibilities: Build high-performance models using distributed...  ...system reliability across our Neuron ecosystem Mentor team members...  ...inference stack Technologies: AWS C# Cloud Java... 
    Amazon Web Service
    Senior
    Performance
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 day ago
  • $193.3k - $261.5k

     ...that help customers change the world. AWS Neuron is the complete software stack...  ...and we are seeking a Senior Software Engineer to join our ML Distributed Training team.In this role...  ...for the development, enablement, and performance optimization of large scale ML model... 
    Amazon Web Service
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $193.3k - $261.5k

     ...Amazon builds Amazon Neuron, the software...  ...Inferentia and Trainium ML accelerators. This...  ...and training performance.The Inference Enablement...  ...boundary, our engineers build systematic...  ...high-performance kernels for ML functions,...  ...optimal performance on AWS ML accelerators.... 
    Amazon Web Service
    Senior
    Performance
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $193.3k - $261.5k

     ...to be part of AI revolution? At AWS our vision is to make deep learning...  ...that make it possible. AWS Neuron is the SDK that optimizes the performance of complex ML models executed on AWS Inferentia...  ...workloadsThis role is for a senior software engineer in the Compiler team for AWS... 
    Amazon Web Service
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $193.3k - $261.5k

    The Product: AWS Machine Learning accelerators...  ...delivers best-in-class ML inference performance at the lowest cost in...  ...stack, the AWS Neuron Software Development...  ...disciplines including silicon engineering, hardware design and...  ....You: As a Sr. Machine Learning Compiler... 
    Amazon Web Service
    Senior
    Performance
    Internship
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $193.5k

     ...of compilers or building ML models on accelerators...  ...software that improves the performance, stability, and user experience of the Neuron compiler. We develop...  ...JAX for deployment on AWS Inferentia and Trainium...  ...architects, runtime and OS engineers, scientists, and ML... 
    Amazon Web Service
    Senior
    Performance
    Full time

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 day ago
  • $140k - $215k

     ...Role:You'll work closely with engineering teams to expand test coverage...  ...that ensures reliability and performance as we deploy AI security...  ...testing in cloud environments (AWS/Azure/GCP) Strong debugging skills...  ...Points:Experience testing AI/ML systems, LLM applications, or... 
    Amazon Web Service
    Senior
    Performance
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  • $184k - $287.5k

     ...skilled and motivated software engineers to join us and build AI...  ...and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry...  ...for the field of ML Systems; survey recent publications...  ...with cloud platforms (AWS/GCP/Azure),... 
    Amazon Web Service
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $183k - $247.6k

    AWS Utility Computing (UC) provides product innovations — from foundational...  ...are seeking a Hardware Design Engineer with role in the definition,...  ...of AWS next generation ML Chips, Cards and server...  ...ways to improve your products performance, quality and cost. We’re changing... 
    Amazon Web Service
    Senior
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $183k - $247.6k

     ...an ongoing basisAWS Compute & ML Services owns the design, planning...  ..., and operation of all AWS global infrastructure. In other...  ...software, hardware, and network engineers, supply chain specialists, security...  ...are industry-leading in performance, frugality and operational excellence... 
    Amazon Web Service
    Senior
    Performance
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    3 days ago
  • $183k - $247.6k

     ...improvements in silicon yield & performance - it’s still Day One here at...  ...experienced Design Verification Engineers to build the next generation...  ...teamInclusive Team CultureHere at AWS, we embrace our differences....  ...complex CPU, GPU, or ML accelerator designsAmazon is an... 
    Amazon Web Service
    Senior
    Performance
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $193.3k - $261.5k

    Annapurna Labs is an integral part of AWS and develops hardware and...  ...AWS customer experience.The AWS Neuron Collectives team is seeking a Software Engineer to optimize collective operations...  ...hardware team, you'll push for maximum performance using C/C++, interfacing with DMA... 
    Amazon Web Service
    Senior
    Performance
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  • $193.5k

     ...collective algorithms and topology choices to maximize training performance Use Neuron Explorer and similar tools to pinpoint bottlenecks in...  ...frameworks and automation solutions Technologies: AI AWS EC2 Firmware Hardware Support Cloud Flow... 
    Amazon Web Service
    Senior
    Performance
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 day ago
  • $193.3k - $261.5k

     ...collective algorithms and topologies to improve training performance. We use tools such as Neuron Explorer to pinpoint bottlenecks in compute and bus...  ...frameworks and automation solutions. Technologies: AI AWS EC2 Firmware Hardware Support Model... 
    Amazon Web Service
    Senior
    Performance
    Full time

    Annapurna Labs Inc.

    Cupertino, CA
    1 day ago
  • $206.9k - $279.9k

    AWS Neuron is looking for an experienced Technical Product Manager...  ...product strategy for the Neuron Kernel Interface (NKI), a compiler...  ..., delivering best-in-class ML performance in the cloud. You will lead...  ...contribute to and influence engineering discussions around technology... 
    Amazon Web Service
    Performance
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $240k

     ...integration, regression, and performance test strategies for AI/ML systems using Python, REST...  ...environments such as AWS and Kubernetes to validate...  ...logs, monitoring tools, and engineering best practices.Document test...  ..., Software QA Engineer, Sr. Software QA Engineer, Senior... 
    Amazon Web Service
    Senior
    Performance

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational...  ...trade-offs• Develop and tune high-performance kernels for critical operations where off-the-...  ...hardware backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how... 
    Amazon Web Service
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    4 days ago
  • $193.3k - $261.5k

     ...Chips) are the brains behind AWS’s Machine Learning servers. Our...  ...looking for a Senior SoC Modeling Engineer to join the team and deliver...  ...our customers.As part of the ML accelerator modeling team, you...  ...and modeling infrastructure performance improvements to help our models... 
    Amazon Web Service
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $212.7k - $287.7k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...for leading a strong team of engineers and managers to help design...  ...libraries with a focus on performance of latest ML models at scale on Trainium...  ...hardware, or experience with CUDA kernels or ML/low-level kernels-... 
    Amazon Web Service
    Performance
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $152k - $241.5k

     ...Artificial Intelligence, High-Performance Computing and Visualization. The...  ...hybrid environment - On‑prem, AWS, GCP, and OCI.Design for...  ...data-driven operations (AIOps/ML-driven signals) that materially...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Amazon Web Service
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $143.7k - $223.6k

     ...hands-on experience with AWS services in production,...  ...: We work alongside engineering peers to develop and maintain high-performance runtime libraries and...  ...build, deploy, and evolve Neuron Runtime and related...  ...enhance the performance of ML kernels and ML frameworks.... 
    Amazon Web Service
    Performance
    Internship
    Flexible hours

    Annapurna Labs Inc.

    Cupertino, CA
    1 day ago
  • $212.7k - $287.7k

     ...We need 2+ years of engineering team management experience...  ...with framework, kernel, runtime, hardware...  ...what unblocks real ML workloads on Trainium...  ...code generation, and performance tuning for custom AWS hardware....  ...NKI team works on AWS Neuron and Trainium, enabling... 
    Amazon Web Service
    Performance
    Full time
    Relocation
    Flexible hours

    Annapurna Labs Inc.

    Cupertino, CA
    1 day ago
  • $165.2k - $223.6k

    AWS Neuron is the complete software stack for the AWS Inferentia and...  ...As the Software Development Engineer for the Neuron Runtime Team,...  ...develop and maintain high-performance runtime libraries and drivers...  ...behavior. Improving performance of ML Kernels and ML Frameworks.In this... 
    Amazon Web Service
    Performance
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $129.3k - $223.6k

     ...understanding of system performance, memory management...  ...in software engineering for large-scale systems...  ...on custom ML hardware accelerators...  ...high-performance kernels and features for ML...  ...operations, utilizing the Neuron architecture and...  ...ML models on AWS accelerators.... 
    Amazon Web Service
    Performance
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 day ago
  • $180k - $220k

     ...are looking for a Senior DevOps Engineer to own the build, deployment, and...  ...systems that push software and ML models to a production fleet of...  ...including networking, storage, performance, and security. ~ Practical experience with AWS and container technologies including... 
    Amazon Web Service
    Senior
    Performance
    Full time

    Knightscope

    Sunnyvale, CA
    16 days ago
  •  ...Design and implement high-performance compute kernels for machine-learning operations using the Neuron architecture and programming...  ...their machine-learning models on AWS accelerators. Collaborate...  ...research and mentor experienced engineers. Requirements At least... 
    Amazon Web Service
    Performance
    Full time
    Internship
    Flexible hours

    Amazon

    Cupertino, CA
    8 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. ML Kernel Performance Engineer, AWS Neuron. Be the first to apply!