Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Software Engineer, AI/ML, AWS Neuron, Model Optimization

$151.3k - $261.5k
Full-time

Annapurna Labs Inc.

Salary: $151,300 - 261,500 per year Requirements:

  • ###
  • Bachelor’s degree in computer science or a related field, or equivalent experience.
  • A minimum of 5 years of professional software development experience.
  • At least 5 years of experience in the design or architecture of new and existing systems, focusing on design patterns, reliability, and scalability.
  • Fundamental knowledge of machine learning and large language models (LLMs), including their architecture, training, and inference lifecycle, alongside experience with optimizations for enhancing model execution.
  • Proficient in software development using C++ and Python (experience in at least one is mandatory).
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Skilled in debugging, profiling, and applying best software engineering practices in large-scale systems.
  • ###
Responsibilities:
  • I will lead efforts in developing distributed inference support for PyTorch within the Neuron SDK.
  • I am responsible for tuning models to achieve the highest performance and efficiency on customer AWS Trainium and Inferentia silicon and servers.
  • My role includes designing, developing, and optimizing machine learning models and frameworks for deployment on specialized ML hardware accelerators.
  • I will participate in all phases of the ML system development lifecycle, encompassing architecture design, implementation, performance profiling, optimizations, testing, and production deployment.
  • I will build infrastructure to systematically analyze and onboard a variety of models with diverse architecture.
  • I am tasked with designing and implementing high-performance kernels and features for ML operations, optimizing system-level performance across various generations of Neuron hardware.
  • I will conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks and implement necessary optimizations including fusion, sharding, and scheduling.
  • My responsibilities also include rigorous testing, both unit and end-to-end model testing, alongside ensuring continuous deployment and releases through pipelines.
  • I will directly collaborate with customers to enable and optimize their ML models on AWS accelerators and work with cross-functional teams to develop innovative optimization techniques.
  • ###
Technologies:
  • AI
  • AWS
  • Hardware
  • Support
  • Machine Learning
  • PyTorch
  • Python
  • Web
  • Cloud
  • Architect
  • Backbone
  • CUDA
  • GitHub
  • LLM

More:

Our team at Annapurna Labs within Amazon Web Services (AWS) focuses on creating AWS Neuron, a software development kit designed to accelerate deep learning and Generative AI workloads. The Neuron SDK is essential for enhancing ML performance on Amazon’s custom machine learning accelerators: Inferentia and Trainium. We pride ourselves on working across all technology layers, from frameworks to hardware, and actively engage with customers to ensure that their workloads run efficiently.

In our unique and experimental work culture, we foster collaboration, technical ownership, and continuous learning, and prioritize mentorship for our newer team members. I am excited to offer a pioneering role at the intersection of machine learning, high-performance computing, and distributed architectures, where I will drive architectural innovations that shape the future of AI acceleration technology.

Join us at the forefront of AI/ML infrastructure challenges and contribute to building impactful solutions for our global customer base!

last updated 37 week of 2026

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Software Engineer, AI/ML, AWS Neuron, Model Optimization in Cupertino, CA vacancy
  • $151.3k - $261.5k

     ...experience in software development...  ...language models (LLMs),...  ...experience in model optimization ~...  ...in software engineering in large-scale...  ..., I will lead efforts to...  ...within the Neuron SDK. I will...  ...efficiency on AWS Trainium...  ...on custom ML hardware accelerators...  ...: AI AWS... 
    Amazon Web Service
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $193.3k - $261.5k

     ...builds Amazon Neuron, the software development...  ...and Trainium ML accelerators....  ...wide range of models and supporting...  ..., our engineers build systematic...  ...fine tuned for optimal performance for...  ...possible in AI acceleration....  ...performance on AWS ML...  ...role will help lead the efforts in... 
    Amazon Web Service
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud...  .... Join us to optimize the latest models to run really fast on the...  ...Sr. Software Development Engineer on the Inference Model Enablement...  ...day in the lifeYou'll lead critical technical... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  • $212.7k - $287.7k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud...  ...accelerators. Join us to optimize LLMs to run really fast on...  ...SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize state... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    7 days ago
  • $129.3k - $223.6k

     ...professional software development experience...  ...language models (LLMs),...  ...experience in optimizations for enhanced...  ...in software engineering for large-scale...  ...deployment on custom ML hardware...  ...the Neuron architecture...  ...ML models on AWS accelerators....  ...Technologies: AI AWS Hardware... 
    Amazon Web Service
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $165.2k - $223.6k

     ...professional software development...  ...language models, including...  ...plus hands-on optimization experience...  ...software engineering practices in...  ...Responsibilities: Lead efforts to...  ...within the Neuron SDK Tune...  ...on AWS Trainium...  ...on custom ML accelerators...  ...Technologies: AI AWS... 
    Amazon Web Service
    Full time
    Internship

    Annapurna Labs Inc.

    Cupertino, CA
    2 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview:...  ...applications of AI & ML are bringing humanity...  ...customers. Our AI models and platforms empower...  ...deploy, and support AI software components including...  ...such as AWS Ultraclusters, Huggingface... 
    Amazon Web Service
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    4 days ago
  •  ...Lead AI Engineer Lead AI Engineer to architect, lead, and...  ...LangChain. Design and optimize AI workflows across...  ...years of experience in AI/ML. ~1+ years of...  ...in Python and modern software engineering practices....  ...with Azure AI Foundry, AWS Bedrock, Google Gemini... 
    Amazon Web Service

    Cardinal Integrated

    Santa Clara, CA
    4 days ago
  • $212.7k

     ...PyTorch or JAX software. We need 3+ years of engineering team...  ..., including model training workflows...  ...on GPUs, Neuron, TPU, or other AI acceleration...  ...kernels or ML and low-level...  ...profiling, and optimization for deep...  ...Responsibilities: We lead a team of...  ...: AI AWS Cloud... 
    Amazon Web Service
    Full time
    Flexible hours

    Amazon Development Center U.S., Inc.

    Cupertino, CA
    18 days ago
  • $193.3k - $261.5k

     ...of experience leading design or architecture...  ...years of full software development...  ...or leading an engineering team ~ Masters...  ...large language model fundamentals,...  ...model execution optimization Responsibilities...  ...across our Neuron ecosystem...  ...Technologies: AWS C# Cloud... 
    Amazon Web Service
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $150k

     ...Machine Learning Engineer About the...  ...of Foundation Models: We are a dedicated...  ...generation of AI builders, and...  ...Engineer focused on ML infrastructure...  ...on AWS (e.g., compute,...  ...frameworks). ~ Solid software engineering skills...  ...of cost optimization, security, and... 
    Amazon Web Service
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    4 days ago
  • $139.23k - $163.8k

     ...Job Summary The Lead Engineer (Generative AI) is a senior...  ...in Large Language Models (LLMs), Retrieval...  ...LangChain, LangGraph, AWS Bedrock, and...  ..., and continuous optimization Implement GenAIOps...  ...security 4. Software Engineering & Architecture...  ..., or AI/ML solutions ~2+ years... 
    Amazon Web Service
    Temporary work
    Work experience placement
    Local area
    3 days per week

    U.S. Bank

    Cupertino, CA
    2 days ago
  • $93k - $128k

     ...least 2 years of engineering team management...  ...and compute graph optimization. We need strong software design...  ...We will have you lead, develop, and mentor...  ...what unblocks real ML workloads on Trainium...  ...: AWS Cloud Hardware...  ...scale. Our AWS Neuron and Trainium products... 
    Amazon Web Service
    Full time
    Relocation

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    18 days ago
  • $212.7k - $287.7k

     ...need 2+ years of engineering team management...  ...and compute graph optimization. We need strong software design...  ...Responsibilities: Lead, develop, and mentor...  ...unblocks real ML workloads on Trainium...  ...for custom AWS hardware. Translate...  ...works on AWS Neuron and Trainium,... 
    Amazon Web Service
    Full time
    Relocation
    Flexible hours

    Annapurna Labs Inc.

    Cupertino, CA
    2 days ago
  • $165.6k

     ...of professional software development experience...  ...using the Neuron architecture and programming models Evaluate and...  ...Apply compiler optimizations such as fusion,...  ...learning models on AWS accelerators...  ...Web Cloud AI Architect...  ...most out of AWS ML accelerators. We... 
    Amazon Web Service
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $143.7k - $223.6k

     ...internship professional software development...  ...with AWS services in production...  ...work alongside engineering peers to...  ...applications and AI accelerators....  ...deploy, and evolve Neuron Runtime and related...  ...help customers optimize AI workloads on...  ...performance of ML kernels and ML... 
    Amazon Web Service
    Full time
    Internship
    Flexible hours

    Annapurna Labs Inc.

    Cupertino, CA
    2 days ago
  • $193.5k

     ...professional software development...  ...experience leading design or...  ...tech lead, or engineering team lead...  ...for ML or HPC, such...  ...GPU kernel optimization and GPGPU computing...  ...using the Neuron...  ...programming models Analyze...  ...models on AWS accelerators...  ...Cloud AI Backbone... 
    Amazon Web Service
    Full time
    Internship
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $143.4k - $165.6k

     ...of professional software development experience...  ...experience with AWS services such as...  ...applications and AI accelerators. We lead the design,...  ...and deployment of Neuron Runtime and related...  ...bottlenecks and better optimize AI workloads...  ...the performance of ML kernels and ML... 
    Amazon Web Service
    Full time
    Internship
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $61k - $101k

     ...training or certification in software engineering concepts, along with 5+...  ...a particular focus on ML systems. We need...  ...enterprise-authorized AI-assisted development tools...  ...platforms, such as AWS or GCP. We are looking...  ...API DDD Java Model Serving More: We... 
    Amazon Web Service
    Full time

    J.P. Morgan

    Palo Alto, CA
    6 days ago
  • $165.6k

     ...internship professional software development...  ..., GPU, vector engines, or ML accelerators....  ...deep learning models, and algorithms...  ...that improve Neuron compiler performance...  ...ML models on AWS accelerators...  ...OpenXLA, and MLIR to optimize advanced ML...  ...TensorFlow AI EC2 IoT... 
    Amazon Web Service
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $179.4k - $204.7k

    Lead AI Engineer (AI Foundations) Overview At Capital...  ...of AI & ML are bringing humanity...  ...customers. Our AI models and platforms empower...  ...deploy, and support AI software components...  ...technologies such as AWS Ultraclusters, Huggingface...  ...-of-the-art LLM optimization techniques to... 
    Amazon Web Service
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  • $229.9k - $262.4k

     ...Overview Senior Lead AI Engineer At Capital One...  ...of AI & ML are bringing humanity...  ...customers.  Our AI models and platforms empower...  ...deploy, and support AI software components...  ...technologies such as AWS Ultraclusters, Huggingface...  ...-of-the-art LLM optimization techniques to... 
    Amazon Web Service
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    3 days ago
  • $193.3k - $261.5k

     ...require 5+ years leading the design or...  ...across the full software development...  ...tech lead, or engineering manager, or hands...  ...experience with ML communication libraries...  ...with our AWS EC2 machines....  ...customers, and large AI models. We...  ...technologies that help optimize the AWS... 
    Amazon Web Service
    Full time

    Amazon.com Services LLC

    Cupertino, CA
    4 days ago
  • $193.5k

     ...experience leading the design...  ...features and optimizations. We need...  ...learning models. We prefer...  ...building ML models on accelerators...  ...maintain software that...  ...of the Neuron compiler....  ...deployment on AWS Inferentia...  ...runtime and OS engineers, scientists...  ...AI EC2 GitHub... 
    Amazon Web Service
    Full time

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    3 days ago
  • $253.1k - $342.3k

    Utility Computing (UC)AWS Utility Computing...  ...that builds AWS Neuron, the software stack that runs all the leading AI models on the AWS...  ...users to develop and optimize AI models on...  ...organization and with ML model developers...  ...qualifications- 10+ years of engineering experience- 5+... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (AI Foundations, LLM Core and Agentic...  ...of AI & ML are bringing humanity...  ...customers. Our AI models and platforms empower...  ...deploy, and support AI software components...  ...technologies such as AWS Ultraclusters, Huggingface...  ...-of-the-art LLM optimization techniques to... 
    Amazon Web Service
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    4 days ago
  • $212.7k - $287.7k

    AWS Neuron is the complete software stack for the AWS Inferentia and Trainium...  ...responsible for leading a strong team of engineers and managers to...  ...of latest ML models at scale on Trainium...  ...techniques in performance optimization, accuracy, and...  ..., TPU or other AI acceleration... 
    Amazon Web Service
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  •  ...Java AI Engineer We are seeking talented Java...  ...experience in AI/ML integration. You will...  ...processes. Optimize applications for performance...  ...understanding of software engineering...  ...or integrating ML models into enterprise solutions...  ...cloud platforms (AWS, GCP, or Azure)... 
    Amazon Web Service

    Kaav Inc.

    Sunnyvale, CA
    4 days ago
  •  ...building the best AI systems for heavy...  ...for Backend AI Engineers to design, build,...  ...possible at scale: model-serving pipelines...  ...blend of backend software engineering, ML infrastructure,...  ...Implement and optimize RAG systems, prompt...  ...infrastructure (AWS, GCP, or Azure) and... 
    Amazon Web Service
    Full time

    Nexxa.ai

    Sunnyvale, CA
    16 hours ago
  • $193.3k - $261.5k

     ...designs silicon and software that accelerates...  ...change the world. AWS Neuron is the complete software...  ...a Senior Software Engineer to join our ML Distributed...  ..., and performance optimization of large scale ML model training across diverse...  ...- 5+ years of leading design or architecture... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Software Engineer, AI/ML, AWS Neuron, Model Optimization. Be the first to apply!