Software Development Manager, LLM Inference Model Enablement, Neuron SDK
$212.7k - $287.7kAmazon Locker
We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scalemachine learning accelerators. Join us to optimize LLMs to run really fast on the Trainium hardware.As an SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.The ideal candidate will have a strong background in LLM model architectures, model performance optimizations, and inference techniques, such as delivering high-performance models using distributed inference libraries. You should be capable of managing demanding, fast-changing priorities. You should have a strong technical ability to understand and deliver as part of a vertically integrated system stack consisting of the PyTorch inference library, Neuron compiler, runtime, and collectives.Key job responsibilities* Management and execution against project plans and delivery commitments* Manage the day-to-day activities of the engineering team* Management of resources, staffing, mentoring, and maintaining a best-of-class engineering team* Report on status of development, quality, operations, and model performance to managementA day in the lifeYou will work with your senior management and technical leaders to define the model enablement and performance optimization for the latest SOTA LLMs, build and deliver them to customers.Meanwhile, lead the team to continue improving the model onboarding experience, as well as enhancing inference usability and quality for Neuron-supported models.You will manage changing priorities as new models and new technologies emerge, and you adapt your team’s work to manage them. You will dive deep to help your team solve technical challenges.Basic qualifications- 3+ years of engineering team management experience- 7+ years of working directly within engineering teams experience- 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- Experience partnering with product or program management teamsPreferred qualification - Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy- Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineersAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 212,700.00 - 287,700.00 USD annually
$193.3k - $261.5k
We develop AWS Neuron, the complete software stack for Trainium, Amazon'... ...optimize the latest models to run really fast on... ...hardware.As a Software Development Engineer on the Inference Model Enablement team, you will... ...engineers, and product managers to architect and deliver...SuggestedInternshipLocal areaFlexible hours$193.3k - $261.5k
...AWS) builds AWS Neuron, the software development kit used to... ...The AWS Neuron SDK, developed by... ...PyTorch and JAX enabling unparalleled ML inference and training... ...wide range of models and supporting... ...wide variety of LLM model families... ..., and product managers to deliver...SuggestedWork experience placementInternshipLocal areaFlexible hours$212.7k - $287.7k
AWS Neuron is the complete software stack for the AWS Inferentia... ...of Software Development for the... ...engineers and managers to help design... ...of latest ML models at scale on... ...engineers focused on enabling new ML... ...on the Neuron SDK / Trainium... ...training and inference solutions. This...SuggestedLocal areaWork from homeFlexible hours$184k - $287.5k
...it. The TensorRT inference platform is the... ...edge deep learning models on every NVIDIA... ...Engineering Manager to take the lead... ...next generation of LLM/VLM/VLA inference software technologies... ...specialized kernel development, runtime optimizations... ...APIs and enabling developers in the...SuggestedFull time$212.7k - $287.7k
...for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler,... ...a talented Software Development Manager with strong leadership and mentoring... ...to market. As deep learning models become more versatile, using compiler...SuggestedLocal areaWork from homeRelocationFlexible hoursDay shift$212.7k - $287.7k
...delivers best-in-class ML inference performance at the... ...cloud. This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML... ...SW Engineering Manager with strong leadership... ...market. As deep learning models become more versatile...Local areaWork from homeRelocationFlexible hours- ...seeking a Senior Product Manager to drive strategy and... ...AMD’s open-source GPU software stack, with a specific focus on large-scale model inference on AMD Instinct™ and... ...and the frameworks that enable it. You will influence... ...understanding of LLM inference, including attention...Remote work
$172.51k - $269.55k
...and laboratory questions. The Software & Informatics Division is one... ...Software Modernization Enablement Lead Engineer to drive the digital... ...paced, next-generation product development environment. You will... ...application modernization.AI/LLM Integration: Hands-on experience...Full timeTemporary workLocal area$117.7k - $221.4k
...and robotics depends not only on stronger models, but also on better infrastructure for... ...the data processing, featurization, and inference foundations that power scalable world understanding... ...{or other frequency dictated by your manager}. This job may be eligible for...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$192.2k - $260k
...'s Delivery Foundation Model team, where you'll work... ...foundation models that enable delivery of billions of... ...amounts of Amazon data and infer at Amazon scale, taking... ...of foundation model development, from multimodal... ...judgment, effectively manage stress and work safely...Local areaWorldwideFlexible hours$250k - $344.5k
...and customer-facing AI-enabled solutions designed to... ...specialized group of software engineers dedicated to... ...Products, you will lead the development of a diverse portfolio... ...and the nuances of LLM integration, while... ...developing top-tier talent and managers focused on the...$174.72k - $295.68k
...Engineer / Research Scientist to drive the modeling and algorithmic development of XPENG’s next-generation Vision-... ...shape the intelligence that enables XPENG’s future L3/L4 autonomous driving... ...foundation or end-to-end driving models, or LLM/VLM architectures (e.g., ViT,...Full time$152k - $241.5k
...universes to explore, enables amazing creativity... ...PyTorch, TRT-LLM, vLLM, SGLang, JAX... ...up to 100K GPUs to inference down at microsecond... ...contribute to the development of innovative... ...on the latest AI models.Improve AI compilers... ...experience) with 5+ software engineering and HPC...Full timeRemote work$224k - $356.5k
...We are seeking a deeply technical software manager to lead production AI inference for NVIDIA Inference Microservices... ...environments. NIM makes state-of-the-art AI models available as production-ready... ...for shipping production‑ready LLM NIMs, including planning, new model...$193.13k - $257.5k
...shipping systems that manage multi-step... ...research, design, development, and deployment of... ...edge deep learning models across all Eightfold... ...a track record of software artifacts or academic... ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills &...Work experience placementWork at officeRemote workFlexible hours3 days per week$212.7k - $287.7k
...customers while enabling delightful experiences... ...-device identity management with privacy-... ...teams measuring model quality, safety,... ...teams on on-device inference, model compression... ...• People Development: Hire, calibrate,... ...experience- 3+ years of Software Engineer, Software...Local areaFlexible hours- ...team in: research, design, development, and deployment of... ...integrate large language models (LLMs) and other state-of... ...proven by a track record of software artifacts or academic... ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in...Full timeWork experience placement
$301.75k - $355k
...Senior Director for the Model LifeCycle team will... ...team and a comprehensive managed platform to oversee the entire application development lifecycle, with a particular... ...preprints in the LLM post-training space. Proficiency... ...on GPU systems and inference frameworks. Benefits...Temporary work$229.9k - $262.4k
...Lead AI Engineer (LLM Gateway, FM Hosting... ...customers. Our AI models and platforms... ...technical program managers, and product managers... ...deploy, and support AI software components... ...large language model inference, similarity search... ...software, and AI enable you to see and exploit...Full timePart timeLocal area$224k - $356.5k
...work directly with software solution providers... ...Omniverse and AI enabled-solutions.Develop... ...innovation and the development of next-generation... ...on AV End-to-End models and GenAI model development... ...in deploying LLM models at scale on... ...and optimize inference latency and throughput...Full time$272k - $431.25k
...generation of interactive world-model systems. With this release,... ...for fidelity, real-time inference performance in world models.... ...Success means more than shipping software: it means high-quality solutions... ..., CI/CD, and release management.Partner with research, simulation...Full time$163k - $236k
...(PT) burndown rates for new model launches to reflect actual capacity... ...of experience in product management or a related technical role.... ...into steps that drive product development.One of the many reasons... ...conversational AI tool that enables users to collaborate with generative...$39 - $66 per hour
...in Silicon Valley focuses on Foundation Models, Big Data Visual Analytics, Explainable AI... ...Autonomous Systems group is responsible for enabling future autonomous Bosch products by... ...position, job location, etc. Your Hiring Manager can share more details about the specific...Work experience placementInternshipLocal areaWorldwide$224k - $356.5k
...s Networking Systems & Software Architecture group is solving... ...and memory management libraries for distributed... ...architectures, KV cache mechanics, model parallelism, or distributed training and inference patterns.Ways to stand... ...vLLM, SGLang, TensorRT-LLM) and their...Full time$87.95k - $203.95k
...apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid... ...Integration Engineer will work with our AI team, software engineers, and business stakeholders to... ..., evaluating metricsHands-On React development lifecycle and component lifecycle...Temporary workWork at officeRemote workFlexible hours$224k - $356.5k
...Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a... ....Work alongside model training, inference, and product divisions to provide trusted... ...experience contributing to open-source software or building platforms, libraries, or tools...Full time$272k - $431.25k
...the future, we’re generating it! Our world model team is pushing the boundaries of... ...AI. We are looking for a Senior Research Manager to lead world-model evaluation and benchmarking... ...evaluations, where access to model internals enables deeper diagnostics, causal analysis, and...Full time$165k - $185k
...Valley focuses on Foundation Models, Big Data Visual Analytics,... ...Foundation Model Powered AI Enablers group plays a pivotal role in... ...expert insights to the management team in relevant technology... ...researchHands-on experience in product development in the above-mentioned areas...Work experience placementWorldwide$246.5k
...to the content they love, enable content publishers to build... ...is our Machine Learning and Inference Platform that powers the entire... ..., design, and lead the development of a SOTA Inference platform... ...that span across hardware, software, and models. We’re looking for a strong...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours- ...Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand,... ...of embodied models Build and manage distributed data processing and ingestion... ...environment; skilled in team development, prioritization, and technical alignment...Full timeWork at officeRemote workWork from homeVisa sponsorshipRelocation packageFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Development Manager, LLM Inference Model Enablement, Neuron SDK. Be the first to apply!
- director of software Cupertino, CA
- software manager Cupertino, CA
- application manager Cupertino, CA
- IT software development manager Cupertino, CA
- software engineer - cloud services Cupertino, CA
- id software Cupertino, CA
- healthcare software sales Cupertino, CA
- software technical support Cupertino, CA
- software asset management analyst Cupertino, CA
- software implementation project manager Cupertino, CA



