Software Development Manager, LLM Inference
$212.7k - $287.7kAnnapurna Labs Inc.
Salary: $212,700 - 287,700 per year Requirements:
- 3 years of experience managing engineering teams
- 7 years of experience working directly in engineering teams
- 3 years of experience designing or architecting new and existing systems, including design patterns, reliability, and scaling
- Experience partnering with product or program management teams
- Strong background in LLM model architectures, model performance optimization, and inference techniques
- Experience delivering high-performance models using distributed inference libraries
- Ability to operate effectively in a fast-changing environment with shifting priorities
- Strong technical understanding of a vertically integrated system stack, including the PyTorch inference library, Neuron compiler, runtime, and collectives
- Experience communicating with users, technical teams, and senior leadership to gather requirements and explain technical designs, product features, and strategy
- Experience recruiting, hiring, mentoring, coaching, and managing software engineering teams to strengthen their effectiveness
- Lead a team of AI/ML engineers to onboard and optimize open-source and customer LLMs, including dense and MoE models, for inference on Trainium accelerators
- Drive improvements in model enablement speed and user experience
- Advance inference usability and quality through features, infrastructure optimization, tools, and automation
- Manage project plans and deliver against commitments
- Oversee the day-to-day activities of the engineering team
- Handle resource planning, staffing, mentoring, and team development
- Report development status, quality, operations, and model performance to management
- Work with senior management and technical leaders to define model enablement and performance optimization for the latest SOTA LLMs
- Build and deliver optimized models to customers
- Help the team solve complex technical challenges
- Adapt team priorities as new models and technologies emerge
- AI
- AWS
- Cloud
- Hardware
- LLM
- Machine Learning
- PyTorch
- Support
More:
We develop AWS Neuron, the complete software stack for Trainium, our custom cloud-scale machine learning accelerators. We are hiring an SDM for the LLM Inference Model Enablement team to lead expert AI/ML engineers focused on optimizing large language models for fast inference on Trainium hardware. You will help improve the onboarding experience for models and enhance inference usability and quality for Neuron-supported models. We offer comprehensive benefits, including health coverage, 401(k) matching, paid time off, parental leave, sign-on payments, and restricted stock units. The role is based in Cupertino, California, with a listed annual base salary range of 212,700 to 287,700 USD.
last updated 37 week of 2026
$212.7k - $287.7k
...develop AWS Neuron, the complete software stack for Trainium, Amazon's... ...hardware.As an SDM for the LLM Inference Model Enablement team, you... .... You should be capable of managing demanding, fast-changing... ...engineering team* Report on status of development, quality, operations, and...SuggestedLocal areaFlexible hours$212.7k - $287.7k
...+ years of experience managing engineering teams ~7... ...requirements, explain software features, outline technical... ..., mentoring, and team development to maintain a top-tier... ...and enhance inference usability and quality... ...Technologies: AWS Cloud LLM Machine Learning...SuggestedFull time- ...AMD is seeking a Senior Product Manager to drive strategy and execution... ...ROCm, AMD’s open-source GPU software stack, with a specific focus on large-scale model inference on AMD Instinct™ and Radeon™ hardware... .... Practical understanding of LLM inference, including attention...SuggestedRemote work
$197.3k - $225.1k
Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible... ...research scientists, technical program managers, and product managers to deliver AI-... ...develop, test, deploy, and support AI software components including foundation model...SuggestedFull timePart timeLocal area$207k - $300k
Work with LLM/Non-LLM models bringup to performance tuning/optimization... ...on Google Cloud TPUs. Manage up to 6 engineers to drive... ..., disaggregated serving, RL inference, etc. Collaborate with the... ....8 years of experience in software development.5 years of experience with one...Suggested- Lead the team in: research, design, development, and deployment of advanced AI... ...innovate, as proven by a track record of software artifacts or academic publications... ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in agentic AI...Full timeWork experience placement
$193.13k - $257.5k
..., shipping systems that manage multi-step reasoning and... ...team in: research, design, development, and deployment of... ...proven by a track record of software artifacts or academic... ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills &...Work experience placementWork at officeRemote workFlexible hours3 days per week$151.3k - $261.5k
...professional experience in software development ~5+ years of experience in... ...architecture, training, and inference lifecycles, with hands-on experience... ...system performance, memory management, and principles of parallel... ...CUDA GitHub LLM Web More: As part of...Full timeInternship$224k - $356.5k
...NVIDIA’s Networking Systems & Software Architecture group is solving... ...communication and memory management libraries for distributed AIDriving... ...or distributed training and inference patterns.Ways to stand out... ...frameworks (vLLM, SGLang, TensorRT-LLM) and their communication...Full time- ...architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud... ...debugging and release decisions, mentor engineers, and write software and automation alongside the team.Release Integration Testing...Work at officeRemote work3 days per week
$87.95k - $203.95k
...apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid... ...Integration Engineer will work with our AI team, software engineers, and business stakeholders to... ..., evaluating metricsHands-On React development lifecycle and component lifecycle...Temporary workWork at officeRemote workFlexible hours$184k - $287.5k
...stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You... ...on scales up to 100K GPUs to inference down at microsecond latency.... ...you ready to contribute to the development of innovative technologies... ...equivalent experience) with 8+ software engineering and HPC/AI...Full timeRemote work$184k - $287.5k
...impact on the world.We are Seeking a Software Engineering Manager for our applied research team within... ...parallelism, or distributed training and inference patterns.Ways to stand out from the... ...frameworks (vLLM, SGLang, TensorRT-LLM) and their communication requirements...Full time$212.7k - $287.7k
AWS Neuron is the complete software stack for the AWS Inferentia... ...them. As the SDM of Software Development for the Neuron Training team... ...strong team of engineers and managers to help design and deploy... ...scale distributed training and inference solutions. This organization...Local areaWork from homeFlexible hours$212.7k - $287.7k
...Inferentia chip delivers best-in-class ML inference performance at the lowest cost in... .... This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML... ...a talented SW Engineering Manager with strong leadership/ mentoring...Local areaWork from homeRelocationFlexible hours$229.9k - $262.4k
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform... ..., technical program managers, and product managers to deliver... ...test, deploy, and support AI software components including foundation... ...and introduce state-of-the-art LLM optimization techniques to improve...Full timePart timeLocal area$197.3k - $225.1k
...AI Engineer (AI Foundations, LLM Core and Agentic AI) Overview... ...scientists, technical program managers, and product managers to... ...test, deploy, and support AI software components including foundation... ...training, large language model inference, similarity search, guardrails...Full timePart timeLocal area$212.7k
...experience working with PyTorch or JAX software. We need 3+ years of engineering team management experience. We need 7+ years... ...scale distributed training and inference solutions as part of the full... ...working practices, and development opportunities. The role is based...Full timeFlexible hours$184k - $287.5k
...technical product marketing manager who is passionate... ...NVIDIA’s AI Platform Software team. We need someone... ...value of training and inference frameworks, such as PyTorch... ...Core, TensorRT LLM and the underlying kernel... ...the use of our software development kits to grow the...Full timeWork experience placement$224k - $356.5k
...you will work directly with software solution providers, developers... ...fostering co-innovation and the development of next-generation solutions.... ...etc.Experience in deploying LLM models at scale on mainstream... ...to profile and optimize inference latency and throughput, memory...Full time$224k - $356.5k
...integration of Generative and Agentic AI software with our highest-priority Enterprise ISV... ...reference architectures that bring RAG, inference, and multi-agent, long-horizon workflows... ...NeMo Agent Toolkit, NIM, Dynamo, TensorRT-LLM) and open tools (vLLM, LangChain, vector...Full time$272k - $431.25k
...in AI: a single RL run ties together inference, rollout, reward and critic evaluation... ...on. We are looking for a Senior Software Engineering Manager to set strategy, build the team, and... ...such as CUDA, NCCL, cuDNN, TensorRT-LLM, Transformer Engine, Nsight, NeMo, or...Full timeRemote work- ...systems as well as internal tooling for system deployment, management and maintenance.What You’ll DoStrategic Technology... ...and NVMe-based solutionsHave hands-on experience with LLM training, fine-tuning, and inference at scaleHave experience with enterprise architecture...Work at officeLocal areaFlexible hours
- ...intelligence, we are a powerhouse - a cutting-edge 'AI Software Solutions Team'. Specialized in AI optimization... ...programming skills in C++ and Python "strong development experience is at least one major DL framework in inference, fine tuning and/or training " MS with years of...
$200k - $300k
...for a Principal AI SoC Runtime Software Architect to own the software... ...role will define and lead development of the end-to-end runtime spanning... ...ingest, preprocessing, AI inference, postprocessing, and delivery... ...heterogeneous workloads, manage ownership, synchronization, and...Contract workFlexible hours$246.5k
...core of this is our Machine Learning and Inference Platform that powers the entire... ...you will architect, design, and lead the development of a SOTA Inference platform that can handle... ...that span across hardware, software, and models. We’re looking for a strong...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours$196k - $310.5k
...(SCG) leads the full product development lifecycle, from early architecture... ...silicon, package, embedded software, testing, and product... ...Own memory performance, power management, thermal, and reliability closure... ...Debug & Root Cause Analysis: Use LLM-assisted tools and ML models...Full timeRemote work$184k - $287.5k
NVIDIA is looking for a brilliant Software & Systems Architect to join the NIC Software/Firmware Architecture group. Be part of a team... ....Explore a wide range of topics, including Generative AI (inference and training), storage, cyber security, HPC, and emulation offloads...Full time$212.7k - $287.7k
Ring is looking for a Software Development Manager responsible for leading a team of engineers working on the development and production of products, focusing on safety and security based connected products, including sensors, networking back-up systems, connected accessories...Local areaWorldwideFlexible hours$212.7k - $287.7k
...concepts are merging the boundaries of control plane and management plane functions and expectedly, software is playing an ever-increasing role in managing... ...enable this vision.We are looking for a software development manager to lead a team of engineers that would build...Local areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Development Manager, LLM Inference. Be the first to apply!
- IT software development manager Cupertino, CA
- software manager Cupertino, CA
- application manager Cupertino, CA
- director of software Cupertino, CA
- internship software Cupertino, CA
- software Cupertino, CA
- software intern Cupertino, CA
- id software Cupertino, CA
- healthcare software sales Cupertino, CA
- entry level software sales Cupertino, CA




