Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Development Manager, LLM Inference

$212.7k - $287.7k
Full-time

Annapurna Labs Inc.

Salary: $212,700 - 287,700 per year Requirements:

  • 3 years of experience managing engineering teams
  • 7 years of experience working directly in engineering teams
  • 3 years of experience designing or architecting new and existing systems, including design patterns, reliability, and scaling
  • Experience partnering with product or program management teams
  • Strong background in LLM model architectures, model performance optimization, and inference techniques
  • Experience delivering high-performance models using distributed inference libraries
  • Ability to operate effectively in a fast-changing environment with shifting priorities
  • Strong technical understanding of a vertically integrated system stack, including the PyTorch inference library, Neuron compiler, runtime, and collectives
  • Experience communicating with users, technical teams, and senior leadership to gather requirements and explain technical designs, product features, and strategy
  • Experience recruiting, hiring, mentoring, coaching, and managing software engineering teams to strengthen their effectiveness
Responsibilities:
  • Lead a team of AI/ML engineers to onboard and optimize open-source and customer LLMs, including dense and MoE models, for inference on Trainium accelerators
  • Drive improvements in model enablement speed and user experience
  • Advance inference usability and quality through features, infrastructure optimization, tools, and automation
  • Manage project plans and deliver against commitments
  • Oversee the day-to-day activities of the engineering team
  • Handle resource planning, staffing, mentoring, and team development
  • Report development status, quality, operations, and model performance to management
  • Work with senior management and technical leaders to define model enablement and performance optimization for the latest SOTA LLMs
  • Build and deliver optimized models to customers
  • Help the team solve complex technical challenges
  • Adapt team priorities as new models and technologies emerge
Technologies:
  • AI
  • AWS
  • Cloud
  • Hardware
  • LLM
  • Machine Learning
  • PyTorch
  • Support

More:

We develop AWS Neuron, the complete software stack for Trainium, our custom cloud-scale machine learning accelerators. We are hiring an SDM for the LLM Inference Model Enablement team to lead expert AI/ML engineers focused on optimizing large language models for fast inference on Trainium hardware. You will help improve the onboarding experience for models and enhance inference usability and quality for Neuron-supported models. We offer comprehensive benefits, including health coverage, 401(k) matching, paid time off, parental leave, sign-on payments, and restricted stock units. The role is based in Cupertino, California, with a listed annual base salary range of 212,700 to 287,700 USD.

last updated 37 week of 2026

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Software Development Manager, LLM Inference in Cupertino, CA vacancy
  • $212.7k - $287.7k

     ...develop AWS Neuron, the complete software stack for Trainium, Amazon's...  ...hardware.As an SDM for the LLM Inference Model Enablement team, you...  .... You should be capable of managing demanding, fast-changing...  ...engineering team* Report on status of development, quality, operations, and... 
    Suggested
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $212.7k - $287.7k

     ...+ years of experience managing engineering teams ~7...  ...requirements, explain software features, outline technical...  ..., mentoring, and team development to maintain a top-tier...  ...and enhance inference usability and quality...  ...Technologies: AWS Cloud LLM Machine Learning... 
    Suggested
    Full time

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    8 days ago
  •  ...AMD is seeking a Senior Product Manager to drive strategy and execution...  ...ROCm, AMD’s open-source GPU software stack, with a specific focus on large-scale model inference on AMD Instinct™ and Radeon™ hardware...  .... Practical understanding of LLM inference, including attention... 
    Suggested
    Remote work

    AMD

    Santa Clara, CA
    4 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible...  ...research scientists, technical program managers, and product managers to deliver AI-...  ...develop, test, deploy, and support AI software components including foundation model... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    4 days ago
  • $207k - $300k

    Work with LLM/Non-LLM models bringup to performance tuning/optimization...  ...on Google Cloud TPUs. Manage up to 6 engineers to drive...  ..., disaggregated serving, RL inference, etc. Collaborate with the...  ....8 years of experience in software development.5 years of experience with one... 
    Suggested

    Google

    Sunnyvale, CA
    3 days ago
  • Lead the team in: research, design, development, and deployment of advanced AI...  ...innovate, as proven by a track record of software artifacts or academic publications...  ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in agentic AI... 
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    9 hours ago
  • $193.13k - $257.5k

     ..., shipping systems that manage multi-step reasoning and...  ...team in: research, design, development, and deployment of...  ...proven by a track record of software artifacts or academic...  ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills &... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    2 days ago
  • $151.3k - $261.5k

     ...professional experience in software development ~5+ years of experience in...  ...architecture, training, and inference lifecycles, with hands-on experience...  ...system performance, memory management, and principles of parallel...  ...CUDA GitHub LLM Web More: As part of... 
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    4 days ago
  • $224k - $356.5k

     ...NVIDIA’s Networking Systems & Software Architecture group is solving...  ...communication and memory management libraries for distributed AIDriving...  ...or distributed training and inference patterns.Ways to stand out...  ...frameworks (vLLM, SGLang, TensorRT-LLM) and their communication... 
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  •  ...architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud...  ...debugging and release decisions, mentor engineers, and write software and automation alongside the team.Release Integration Testing... 
    Work at office
    Remote work
    3 days per week

    Cerebras Systems

    Sunnyvale, CA
    8 hours ago
  • $87.95k - $203.95k

     ...apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid...  ...Integration Engineer will work with our AI team, software engineers, and business stakeholders to...  ..., evaluating metricsHands-On React development lifecycle and component lifecycle... 
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You...  ...on scales up to 100K GPUs to inference down at microsecond latency....  ...you ready to contribute to the development of innovative technologies...  ...equivalent experience) with 8+ software engineering and HPC/AI... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...impact on the world.We are Seeking a Software Engineering Manager for our applied research team within...  ...parallelism, or distributed training and inference patterns.Ways to stand out from the...  ...frameworks (vLLM, SGLang, TensorRT-LLM) and their communication requirements... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $212.7k - $287.7k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...them. As the SDM of Software Development for the Neuron Training team...  ...strong team of engineers and managers to help design and deploy...  ...scale distributed training and inference solutions. This organization... 
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    9 hours ago
  • $212.7k - $287.7k

     ...Inferentia chip delivers best-in-class ML inference performance at the lowest cost in...  .... This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML...  ...a talented SW Engineering Manager with strong leadership/ mentoring... 
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform...  ..., technical program managers, and product managers to deliver...  ...test, deploy, and support AI software components including foundation...  ...and introduce state-of-the-art LLM optimization techniques to improve... 
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    1 day ago
  • $197.3k - $225.1k

     ...AI Engineer (AI Foundations, LLM Core and Agentic AI) Overview...  ...scientists, technical program managers, and product managers to...  ...test, deploy, and support AI software components including foundation...  ...training, large language model inference, similarity search, guardrails... 
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    5 days ago
  • $212.7k

     ...experience working with PyTorch or JAX software. We need 3+ years of engineering team management experience. We need 7+ years...  ...scale distributed training and inference solutions as part of the full...  ...working practices, and development opportunities. The role is based... 
    Full time
    Flexible hours

    Amazon Development Center U.S., Inc.

    Cupertino, CA
    19 days ago
  • $184k - $287.5k

     ...technical product marketing manager who is passionate...  ...NVIDIA’s AI Platform Software team. We need someone...  ...value of training and inference frameworks, such as PyTorch...  ...Core, TensorRT LLM and the underlying kernel...  ...the use of our software development kits to grow the... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...you will work directly with software solution providers, developers...  ...fostering co-innovation and the development of next-generation solutions....  ...etc.Experience in deploying LLM models at scale on mainstream...  ...to profile and optimize inference latency and throughput, memory... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...integration of Generative and Agentic AI software with our highest-priority Enterprise ISV...  ...reference architectures that bring RAG, inference, and multi-agent, long-horizon workflows...  ...NeMo Agent Toolkit, NIM, Dynamo, TensorRT-LLM) and open tools (vLLM, LangChain, vector... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...in AI: a single RL run ties together inference, rollout, reward and critic evaluation...  ...on. We are looking for a Senior Software Engineering Manager to set strategy, build the team, and...  ...such as CUDA, NCCL, cuDNN, TensorRT-LLM, Transformer Engine, Nsight, NeMo, or... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...systems as well as internal tooling for system deployment, management and maintenance.What You’ll DoStrategic Technology...  ...and NVMe-based solutionsHave hands-on experience with LLM training, fine-tuning, and inference at scaleHave experience with enterprise architecture... 
    Work at office
    Local area
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  •  ...intelligence, we are a powerhouse - a cutting-edge 'AI Software Solutions Team'. Specialized in AI optimization...  ...programming skills in C++ and Python   "strong development experience is at least one major DL framework in inference, fine tuning and/or training " MS with years of... 

    AMD

    Santa Clara, CA
    1 day ago
  • $200k - $300k

     ...for a Principal AI SoC Runtime Software Architect to own the software...  ...role will define and lead development of the end-to-end runtime spanning...  ...ingest, preprocessing, AI inference, postprocessing, and delivery...  ...heterogeneous workloads, manage ownership, synchronization, and... 
    Contract work
    Flexible hours

    Velaura

    Santa Clara, CA
    3 days ago
  • $246.5k

     ...core of this is our Machine Learning and Inference Platform that powers the entire...  ...you will architect, design, and lead the development of a SOTA Inference platform that can handle...  ...that span across hardware, software, and models. We’re looking for a strong... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    9 hours ago
  • $196k - $310.5k

     ...(SCG) leads the full product development lifecycle, from early architecture...  ...silicon, package, embedded software, testing, and product...  ...Own memory performance, power management, thermal, and reliability closure...  ...Debug & Root Cause Analysis: Use LLM-assisted tools and ML models... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    NVIDIA is looking for a brilliant Software & Systems Architect to join the NIC Software/Firmware Architecture group. Be part of a team...  ....Explore a wide range of topics, including Generative AI (inference and training), storage, cyber security, HPC, and emulation offloads... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $212.7k - $287.7k

    Ring is looking for a Software Development Manager responsible for leading a team of engineers working on the development and production of products, focusing on safety and security based connected products, including sensors, networking back-up systems, connected accessories... 
    Local area
    Worldwide
    Flexible hours

    Amazon

    Sunnyvale, CA
    1 day ago
  • $212.7k - $287.7k

     ...concepts are merging the boundaries of control plane and management plane functions and expectedly, software is playing an ever-increasing role in managing...  ...enable this vision.We are looking for a software development manager to lead a team of engineers that would build... 
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Development Manager, LLM Inference. Be the first to apply!