Sr. ML Kernel Performance Engineer, AWS Neuron
$193.5kAnnapurna Labs (U.S.) Inc.
Salary: $193,500 - 193,500 per year Requirements:
- 5+ years of non-internship professional software development experience
- 5+ years of programming experience in at least one software language
- 5+ years of experience leading design or architecture for new and existing systems, including design patterns, reliability, and scaling
- 5+ years of full software development life cycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations
- Experience serving as a mentor, tech lead, or engineering team lead
- Bachelors degree in computer science or an equivalent background
- 6+ years of full software development experience
- Expertise in accelerator architectures for ML or HPC, such as GPUs, CPUs, FPGAs, or custom architectures
- Experience with GPU kernel optimization and GPGPU computing, such as CUDA, NKI, Triton, OpenCL, SYCL, or ROCm
- Demonstrated experience with NVIDIA PTX and/or AMD GPU ISA
- Experience building high-performance libraries for HPC applications
- Proficiency in low-level performance tuning for GPUs
- Experience developing LLVM/MLIR backends for GPUs
- Knowledge of ML frameworks such as PyTorch or TensorFlow and their GPU backends
- Experience with parallel programming and optimization methods
- Understanding of GPU memory hierarchies and optimization strategies
- Design and build high-performance compute kernels for ML operations using the Neuron architecture and programming models
- Analyze and improve kernel-level performance across multiple generations of Neuron hardware
- Use profiling tools to perform detailed performance analysis and identify bottlenecks
- Develop compiler optimizations including fusion, sharding, tiling, and scheduling
- Partner directly with customers to enable and optimize their ML models on AWS accelerators
- Work across teams to create innovative kernel optimization techniques
- Architect and implement business-critical features
- Publish cutting-edge research
- Mentor experienced engineers
- AWS
- Architect
- CUDA
- Hardware
- LLVM
- Machine Learning
- PyTorch
- TensorFlow
- Web
- Cloud
- AI
- Backbone
- Backend
- Flow
- Support
More:
We are the Annapurna Labs team at Amazon Web Services, building AWS Neuron, the software development kit that accelerates deep learning and GenAI workloads on our custom machine learning accelerators, Inferentia and Trainium. Our Acceleration Kernel Library team focuses on maximizing performance at the hardware-software boundary, crafting high-performance kernels for ML functions so every FLOP counts for customer workloads. We work across the Neuron Compiler organization, spanning frameworks, compilers, runtime, and collectives, while also contributing to future architecture designs and collaborating closely with customers on model enablement. This role offers the chance to work on cutting-edge products at the intersection of machine learning, high-performance computing, and distributed architectures. We value an innovative, agile, and collaborative culture, and we offer flexible working arrangements, a strong work-life balance, inclusive team programs, and comprehensive benefits including health coverage, retirement matching, paid time off, parental leave, and sign-on payments and RSUs. The position is based in Cupertino, California, with a listed annual base salary of 193,500.00 USD.
last updated 37 week of 2026
$193.3k - $261.5k
...at Amazon Web Services (AWS) builds AWS Neuron, the software development... ...Trainium.The Acceleration Kernel Library team is at the forefront of maximizing performance for AWS's custom ML accelerators. Working at... ...hardware-software boundary, our engineers craft high-performance...Amazon Web ServiceSeniorPerformanceInternshipLocal areaWork from homeFlexible hours$165.6k
...Responsibilities: Design and build high-performance compute kernels for machine learning operations using the Neuron architecture and programming... ...machine learning models on AWS accelerators Partner with... ...get the most out of AWS ML accelerators. We operate across...Amazon Web ServicePerformanceFull timeInternship$253.1k - $342.3k
Utility Computing (UC)AWS Utility Computing (... ...that builds AWS Neuron, the software stack... ...accelerators.As a Sr. Software Development... ...and with ML model developers and... ...SDK to deliver best performance and usability on top... ...qualifications- 10+ years of engineering experience- 5+...Amazon Web ServiceSeniorPerformanceLocal areaFlexible hours$193.3k - $261.5k
We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom... ...fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team... ...job responsibilities* Deliver high-performance models using distributed inference...Amazon Web ServiceSeniorPerformanceInternshipLocal areaFlexible hours$193.3k - $261.5k
...as a tech lead, or leading an engineering team ~ Masters degree in... ...Responsibilities: Build high-performance models using distributed... ...system reliability across our Neuron ecosystem Mentor team members... ...inference stack Technologies: AWS C# Cloud Java...Amazon Web ServiceSeniorPerformanceFull timeInternship$193.3k - $261.5k
...that help customers change the world. AWS Neuron is the complete software stack... ...and we are seeking a Senior Software Engineer to join our ML Distributed Training team.In this role... ...for the development, enablement, and performance optimization of large scale ML model...Amazon Web ServiceSeniorPerformanceInternshipLocal areaFlexible hours$193.3k - $261.5k
...Amazon builds Amazon Neuron, the software... ...Inferentia and Trainium ML accelerators. This... ...and training performance.The Inference Enablement... ...boundary, our engineers build systematic... ...high-performance kernels for ML functions,... ...optimal performance on AWS ML accelerators....Amazon Web ServiceSeniorPerformanceWork experience placementInternshipLocal areaFlexible hours$193.3k - $261.5k
...to be part of AI revolution? At AWS our vision is to make deep learning... ...that make it possible. AWS Neuron is the SDK that optimizes the performance of complex ML models executed on AWS Inferentia... ...workloadsThis role is for a senior software engineer in the Compiler team for AWS...Amazon Web ServiceSeniorPerformanceLocal areaFlexible hours$193.3k - $261.5k
The Product: AWS Machine Learning accelerators... ...delivers best-in-class ML inference performance at the lowest cost in... ...stack, the AWS Neuron Software Development... ...disciplines including silicon engineering, hardware design and... ....You: As a Sr. Machine Learning Compiler...Amazon Web ServiceSeniorPerformanceInternshipLocal areaWork from homeRelocationFlexible hours$193.5k
...of compilers or building ML models on accelerators... ...software that improves the performance, stability, and user experience of the Neuron compiler. We develop... ...JAX for deployment on AWS Inferentia and Trainium... ...architects, runtime and OS engineers, scientists, and ML...Amazon Web ServiceSeniorPerformanceFull time$140k - $215k
...Role:You'll work closely with engineering teams to expand test coverage... ...that ensures reliability and performance as we deploy AI security... ...testing in cloud environments (AWS/Azure/GCP) Strong debugging skills... ...Points:Experience testing AI/ML systems, LLM applications, or...Amazon Web ServiceSeniorPerformanceFull timeContract workWork experience placementWork at officeLocal area$184k - $287.5k
...skilled and motivated software engineers to join us and build AI... ...and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... ...for the field of ML Systems; survey recent publications... ...with cloud platforms (AWS/GCP/Azure),...Amazon Web ServiceSeniorPerformanceFull time$183k - $247.6k
AWS Utility Computing (UC) provides product innovations — from foundational... ...are seeking a Hardware Design Engineer with role in the definition,... ...of AWS next generation ML Chips, Cards and server... ...ways to improve your products performance, quality and cost. We’re changing...Amazon Web ServiceSeniorPerformanceLocal areaFlexible hours$183k - $247.6k
...an ongoing basisAWS Compute & ML Services owns the design, planning... ..., and operation of all AWS global infrastructure. In other... ...software, hardware, and network engineers, supply chain specialists, security... ...are industry-leading in performance, frugality and operational excellence...Amazon Web ServiceSeniorPerformanceLocal areaFlexible hours$183k - $247.6k
...improvements in silicon yield & performance - it’s still Day One here at... ...experienced Design Verification Engineers to build the next generation... ...teamInclusive Team CultureHere at AWS, we embrace our differences.... ...complex CPU, GPU, or ML accelerator designsAmazon is an...Amazon Web ServiceSeniorPerformanceLocal areaWork from homeFlexible hours$193.3k - $261.5k
Annapurna Labs is an integral part of AWS and develops hardware and... ...AWS customer experience.The AWS Neuron Collectives team is seeking a Software Engineer to optimize collective operations... ...hardware team, you'll push for maximum performance using C/C++, interfacing with DMA...Amazon Web ServiceSeniorPerformanceLocal areaWork from homeFlexible hours$193.5k
...collective algorithms and topology choices to maximize training performance Use Neuron Explorer and similar tools to pinpoint bottlenecks in... ...frameworks and automation solutions Technologies: AI AWS EC2 Firmware Hardware Support Cloud Flow...Amazon Web ServiceSeniorPerformanceFlexible hours$193.3k - $261.5k
...collective algorithms and topologies to improve training performance. We use tools such as Neuron Explorer to pinpoint bottlenecks in compute and bus... ...frameworks and automation solutions. Technologies: AI AWS EC2 Firmware Hardware Support Model...Amazon Web ServiceSeniorPerformanceFull time$206.9k - $279.9k
AWS Neuron is looking for an experienced Technical Product Manager... ...product strategy for the Neuron Kernel Interface (NKI), a compiler... ..., delivering best-in-class ML performance in the cloud. You will lead... ...contribute to and influence engineering discussions around technology...Amazon Web ServicePerformanceFlexible hours$240k
...integration, regression, and performance test strategies for AI/ML systems using Python, REST... ...environments such as AWS and Kubernetes to validate... ...logs, monitoring tools, and engineering best practices.Document test... ..., Software QA Engineer, Sr. Software QA Engineer, Senior...Amazon Web ServiceSeniorPerformance$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational... ...trade-offs• Develop and tune high-performance kernels for critical operations where off-the-... ...hardware backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how...Amazon Web ServiceSeniorPerformanceInternshipLocal areaFlexible hours$193.3k - $261.5k
...Chips) are the brains behind AWS’s Machine Learning servers. Our... ...looking for a Senior SoC Modeling Engineer to join the team and deliver... ...our customers.As part of the ML accelerator modeling team, you... ...and modeling infrastructure performance improvements to help our models...Amazon Web ServiceSeniorPerformanceInternshipLocal areaFlexible hours$212.7k - $287.7k
AWS Neuron is the complete software stack for the AWS Inferentia... ...for leading a strong team of engineers and managers to help design... ...libraries with a focus on performance of latest ML models at scale on Trainium... ...hardware, or experience with CUDA kernels or ML/low-level kernels-...Amazon Web ServicePerformanceLocal areaWork from homeFlexible hours$152k - $241.5k
...Artificial Intelligence, High-Performance Computing and Visualization. The... ...hybrid environment - On‑prem, AWS, GCP, and OCI.Design for... ...data-driven operations (AIOps/ML-driven signals) that materially... ...Perl, or Ruby.Mentored other engineers and influenced technical direction...Amazon Web ServiceSeniorPerformanceFull time$143.7k - $223.6k
...hands-on experience with AWS services in production,... ...: We work alongside engineering peers to develop and maintain high-performance runtime libraries and... ...build, deploy, and evolve Neuron Runtime and related... ...enhance the performance of ML kernels and ML frameworks....Amazon Web ServicePerformanceInternshipFlexible hours$212.7k - $287.7k
...We need 2+ years of engineering team management experience... ...with framework, kernel, runtime, hardware... ...what unblocks real ML workloads on Trainium... ...code generation, and performance tuning for custom AWS hardware.... ...NKI team works on AWS Neuron and Trainium, enabling...Amazon Web ServicePerformanceFull timeRelocationFlexible hours$165.2k - $223.6k
AWS Neuron is the complete software stack for the AWS Inferentia and... ...As the Software Development Engineer for the Neuron Runtime Team,... ...develop and maintain high-performance runtime libraries and drivers... ...behavior. Improving performance of ML Kernels and ML Frameworks.In this...Amazon Web ServicePerformanceInternshipLocal areaWork from homeFlexible hours$129.3k - $223.6k
...understanding of system performance, memory management... ...in software engineering for large-scale systems... ...on custom ML hardware accelerators... ...high-performance kernels and features for ML... ...operations, utilizing the Neuron architecture and... ...ML models on AWS accelerators....Amazon Web ServicePerformanceFull timeInternship$180k - $220k
...are looking for a Senior DevOps Engineer to own the build, deployment, and... ...systems that push software and ML models to a production fleet of... ...including networking, storage, performance, and security. ~ Practical experience with AWS and container technologies including...Amazon Web ServiceSeniorPerformanceFull time- ...Design and implement high-performance compute kernels for machine-learning operations using the Neuron architecture and programming... ...their machine-learning models on AWS accelerators. Collaborate... ...research and mentor experienced engineers. Requirements At least...Amazon Web ServicePerformanceFull timeInternshipFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. ML Kernel Performance Engineer, AWS Neuron. Be the first to apply!
- machine learning engineer Cupertino, CA
- computer vision machine learning engineer Cupertino, CA
- senior ml engineer Cupertino, CA
- senior developer Cupertino, CA
- senior aws cloud engineer Cupertino, CA
- remote senior salesforce administrator Cupertino, CA
- senior manager tax Cupertino, CA
- senior tax Cupertino, CA
- senior helpdesk technician Cupertino, CA
- senior principal cloud computing engineer Cupertino, CA

