Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Development Manager, LLM Inference Model Enablement, Neuron SDK

$212.7k - $287.7k

Annapurna Labs (U.S.)

We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale

machine learning accelerators. Join us to optimize LLMs to run really fast on the Trainium hardware.

As an SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.

The ideal candidate will have a strong background in LLM model architectures, model performance optimizations, and inference techniques, such as delivering high-performance models using distributed inference libraries. You should be capable of managing demanding, fast-changing priorities. You should have a strong technical ability to understand and deliver as part of a vertically integrated system stack consisting of the PyTorch inference library, Neuron compiler, runtime, and collectives.

Key job responsibilities

  • Management and execution against project plans and delivery commitments
  • Manage the day-to-day activities of the engineering team
  • Management of resources, staffing, mentoring, and maintaining a best-of-class engineering team
  • Report on status of development, quality, operations, and model performance to management

A day in the life

You will work with your senior management and technical leaders to define the model enablement and performance optimization for the latest SOTA LLMs, build and deliver them to customers.

Meanwhile, lead the team to continue improving the model onboarding experience, as well as enhancing inference usability and quality for Neuron-supported models.

You will manage changing priorities as new models and new technologies emerge, and you adapt your team's work to manage them. You will dive deep to help your team solve technical challenges.

BASIC QUALIFICATIONS

  • 3+ years of engineering team management experience
  • 7+ years of working directly within engineering teams experience
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • Experience partnering with product or program management teams

PREFERRED QUALIFICATIONS

  • Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy
  • Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company's reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at

USA, CA, Cupertino - 212,700.00 - 287,700.00 USD annually

Vacancy posted 9 days ago
Similar jobs that could be interesting for youBased on the Software Development Manager, LLM Inference Model Enablement, Neuron SDK in Cupertino, CA vacancy
  • $212.7k - $287.7k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom...  ....As an SDM for the LLM Inference Model Enablement team, you will lead a team...  ...should be capable of managing demanding, fast-...  ...team* Report on status of development, quality, operations, and... 
    Suggested
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $212.7k - $287.7k

     ...experience managing engineering...  ...requirements, explain software features,...  ..., and team development to maintain...  ..., and model performance...  ...define model enablement and performance...  ...and enhance inference usability and...  ...quality for Neuron-supported...  ...Cloud LLM Machine Learning... 
    Suggested
    Full time

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    9 days ago
  • $193.3k - $261.5k

     ...builds Amazon Neuron, the software development kit used to accelerate...  ...Amazon Neuron SDK, developed by...  ...and JAX enabling unparalleled ML inference and training...  ...wide range of models and supporting...  ...wide variety of LLM model families...  ..., and product managers to deliver... 
    Suggested
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon'...  ...optimize the latest models to run really fast on...  ...hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will...  ...engineers, and product managers to architect and deliver... 
    Suggested
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $212.7k - $287.7k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...of Software Development for the...  ...engineers and managers to help design...  ...of latest ML models at scale on...  ...engineers focused on enabling new ML...  ...on the Neuron SDK / Trainium...  ...training and inference solutions. This... 
    Suggested
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $212.7k - $287.7k

     ...of experience managing engineering teams...  ...background in LLM model architectures,...  ..., and inference techniques ~...  ...inference library, Neuron compiler, runtime...  ...and managing software engineering...  ...improvements in model enablement speed and user...  ..., and team development Report... 
    Full time
    Shift work

    Annapurna Labs Inc.

    Cupertino, CA
    6 days ago
  • $212.7k

     ...PyTorch or JAX software. We need...  ...team management experience....  ..., including model training workflows...  ...production on GPUs, Neuron, TPU, or...  ...focused on enabling new machine...  ...the Neuron SDK and Trainium...  ...and inference solutions as...  ...practices, and development opportunities... 
    Full time
    Flexible hours

    Amazon Development Center U.S., Inc.

    Cupertino, CA
    20 days ago
  • $151.3k - $261.5k

     ...professional software development experience....  ...large language models (LLMs), including...  ...training, and inference lifecycle, alongside...  ..., memory management, and parallel...  ...PyTorch within the Neuron SDK. ~ I am...  ...customers to enable and optimize their...  ...GitHub LLM More: Our... 
    Full time

    Annapurna Labs Inc.

    Cupertino, CA
    4 days ago
  • $151.3k - $261.5k

     ...experience in software development ~5+ years of...  ...large language models (LLMs),...  ...training, and inference lifecycles, with...  ...performance, memory management, and...  ...PyTorch within the Neuron SDK. I will optimize...  ...customers to enable and optimize their...  ...GitHub LLM Web More... 
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    5 days ago
  •  ...leading training and inference speeds; over 10...  ...the leading model labs, global...  ...Model Scaling team enables state-of-the-...  ...as well as development of high-performance...  ..., product management, and AI research...  ...support for emerging LLM architectures...  ...hardware/software co-design through... 

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $212.7k - $287.7k

     ...for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler,...  ...a talented Software Development Manager with strong leadership and mentoring...  ...to market. As deep learning models become more versatile, using compiler... 
    Local area
    Work from home
    Relocation
    Flexible hours
    Day shift

    Amazon

    Cupertino, CA
    1 day ago
  • $253.1k - $342.3k

     ..., you’ll support the development and management of Compute, Database,...  ...team that builds AWS Neuron, the software stack that runs all the leading AI models on the AWS Inferentia...  ...Manager in the AWS Neuron SDK organization you will...  ...and libraries that enable users to develop and... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $212.7k - $287.7k

     ...delivers best-in-class ML inference performance at the...  ...cloud. This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML...  ...SW Engineering Manager with strong leadership...  ...market. As deep learning models become more versatile... 
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...seeking a Senior Product Manager to drive strategy and...  ...AMD’s open-source GPU software stack, with a specific focus on large-scale model inference on AMD Instinct™ and...  ...and the frameworks that enable it. You will influence...  ...understanding of LLM inference, including attention... 
    Remote work

    AMD

    Santa Clara, CA
    13 hours ago
  • $229.9k - $262.4k

     ...Engineer (FM Hosting, LLM Inference) Overview: At Capital...  ...of customers. Our AI models and platforms empower...  ..., technical program managers, and product managers...  ...deploy, and support AI software components including foundation...  ..., software, and AI enable you to see and exploit... 
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    2 days ago
  • $212.7k - $287.7k

     ...years of engineering team management experience. We need 6+...  ...optimization. We need strong software design fundamentals and...  ...including hiring, career development, and maintaining team...  ...Our NKI team works on AWS Neuron and Trainium, enabling high-performance machine learning... 
    Full time
    Relocation
    Flexible hours

    Annapurna Labs Inc.

    Cupertino, CA
    4 days ago
  • $192k - $278k

     ...product strategy for AI model serving safety...  ...for AI model inference, including real-time...  ...grade safety controls enabling custom threshold...  ...abuse triage and manage regulatory escalation...  ...Language Model (LLM) interfaces into...  ...relationships, commercial co-development contracts, and... 

    Google

    Sunnyvale, CA
    2 days ago
  • $207k - $300k

    Work with LLM/Non-LLM models bringup to performance tuning/...  ...Google Cloud TPUs. Manage up to 6 engineers to...  ...disaggregated serving, RL inference, etc. Collaborate...  ...of experience in software development.5 years of experience...  ...trusted partner to enable growth and solve their... 

    Google

    Sunnyvale, CA
    4 days ago
  • $93k - $128k

     ...2 years of engineering team management experience. We look for 6...  ...optimization. We need strong software design fundamentals and...  ...customers at cloud scale. Our AWS Neuron and Trainium products power...  ...through the Neuron Software Development Kit, including the Neuron... 
    Full time
    Relocation

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    20 days ago
  • $192.2k - $260k

     ...'s Delivery Foundation Model team, where you'll work...  ...foundation models that enable delivery of billions of...  ...amounts of Amazon data and infer at Amazon scale, taking...  ...of foundation model development, from multimodal...  ...judgment, effectively manage stress and work safely... 
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    13 hours ago
  • $174.72k - $295.68k

     ...Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-...  ...shape the intelligence that enables XPENG’s future L3/L4 autonomous driving...  ...foundation or end-to-end driving models, or LLM/VLM architectures (e.g., ViT,... 
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $136.8k - $277.2k

     ...company\'s risk management, governance and...  ...products that enable and empower continuous...  ...Model Evaluation & Audit...  ...Professional Development: Continue to develop...  ...transformer-based LLM architectures (...  ...pipelines, and inference behaviors....  ...end and back end software development skills... 
    Temporary work
    Local area

    TikTok

    San Jose, CA
    1 day ago
  • $200k - $300k

     ...AI SoC Runtime Software Architect to own...  ...and lead development of the end-to-end...  ...preprocessing, AI inference, postprocessing...  ...heterogeneous workloads, manage ownership,...  ...and SDK components....  ...wide execution model for coordinating...  ...hardware layers and enables optimization against... 
    Contract work
    Flexible hours

    Velaura

    Santa Clara, CA
    4 days ago
  •  ...team in: research, design, development, and deployment of...  ...integrate large language models (LLMs) and other state-of...  ...proven by a track record of software artifacts or academic...  ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in... 
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    13 hours ago
  • $193.13k - $257.5k

     ...shipping systems that manage multi-step...  ...research, design, development, and deployment of...  ...edge deep learning models across all Eightfold...  ...a track record of software artifacts or academic...  ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills &... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...universes to explore, enables amazing creativity...  ...PyTorch, TRT-LLM, vLLM, SGLang, JAX...  ...up to 100K GPUs to inference down at microsecond...  ...contribute to the development of innovative...  ...on the latest AI models.Improve AI compilers...  ...experience) with 8+ software engineering and HPC... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $207k - $300k

    Build infrastructure for model creators to be able to...  ...for partners to manage their services on Vertex...  ...years of experience in software development.5 years of experience...  ...generative AI tools or LLM interfaces into workflows...  ...trusted partner to enable growth and solve their... 

    Google

    Sunnyvale, CA
    2 days ago
  • $272k - $431.25k

     ...generation of interactive world-model systems. With this release,...  ...for fidelity, real-time inference performance in world models....  ...Success means more than shipping software: it means high-quality solutions...  ..., CI/CD, and release management.Partner with research, simulation... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...work directly with software solution providers...  ...Omniverse and AI enabled-solutions.Develop...  ...innovation and the development of next-generation...  ...on AV End-to-End models and GenAI model development...  ...in deploying LLM models at scale on...  ...and optimize inference latency and throughput... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $219k - $351k

     ...bandwidth business. As models scale past what any...  ...the core product of AI inference, not an afterthought.We...  ...Samsung Cognos, AI memory software that moves model state...  ...Dynamo, TensorRT-LLM, llama.cpp-class engines...  ...their memory-management internals, not just their... 
    Work at office
    Remote work
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Development Manager, LLM Inference Model Enablement, Neuron SDK. Be the first to apply!