Software Development Manager, LLM Inference Model Enablement, Neuron SDK
$212.7k - $287.7kAmazon Locker
We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scalemachine learning accelerators. Join us to optimize LLMs to run really fast on the Trainium hardware.As an SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.The ideal candidate will have a strong background in LLM model architectures, model performance optimizations, and inference techniques, such as delivering high-performance models using distributed inference libraries. You should be capable of managing demanding, fast-changing priorities. You should have a strong technical ability to understand and deliver as part of a vertically integrated system stack consisting of the PyTorch inference library, Neuron compiler, runtime, and collectives.Key job responsibilities* Management and execution against project plans and delivery commitments* Manage the day-to-day activities of the engineering team* Management of resources, staffing, mentoring, and maintaining a best-of-class engineering team* Report on status of development, quality, operations, and model performance to managementA day in the lifeYou will work with your senior management and technical leaders to define the model enablement and performance optimization for the latest SOTA LLMs, build and deliver them to customers.Meanwhile, lead the team to continue improving the model onboarding experience, as well as enhancing inference usability and quality for Neuron-supported models.You will manage changing priorities as new models and new technologies emerge, and you adapt your team’s work to manage them. You will dive deep to help your team solve technical challenges.Basic qualifications- 3+ years of engineering team management experience- 7+ years of working directly within engineering teams experience- 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- Experience partnering with product or program management teamsPreferred qualification - Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy- Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineersAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 212,700.00 - 287,700.00 USD annually
$165.2k - $223.6k
...AWS) builds AWS Neuron, the software development kit used to... ...The AWS Neuron SDK, developed by... ...PyTorch and JAX enabling unparalleled ML inference and training... ...wide range of models and supporting... ...wide variety of LLM model families... ..., and product managers to deliver...SuggestedWork experience placementInternshipLocal areaFlexible hours$193.3k - $261.5k
We develop AWS Neuron, the complete software stack for Trainium, Amazon'... ...optimize the latest models to run really fast on... ...hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will... ...engineers, and product managers to architect and deliver...SuggestedInternshipLocal areaFlexible hours$212.7k - $287.7k
AWS Neuron is the complete software stack for the AWS Inferentia... ...of Software Development for the... ...engineers and managers to help design... ...of latest ML models at scale on... ...engineers focused on enabling new ML... ...on the Neuron SDK / Trainium... ...training and inference solutions. This...SuggestedLocal areaWork from homeFlexible hours- ...leading training and inference speeds; over 10... ...the leading model labs, global... ...Model Scaling team enables state-of-the-... ...as well as development of high-performance... ..., product management, and AI research... ...support for emerging LLM architectures... ...hardware/software co-design through...Suggested
$212.7k - $287.7k
...for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler,... ...a talented Software Development Manager with strong leadership and mentoring... ...to market. As deep learning models become more versatile, using compiler...SuggestedLocal areaWork from homeRelocationFlexible hoursDay shift$212.7k - $287.7k
...delivers best-in-class ML inference performance at the... ...cloud. This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML... ...SW Engineering Manager with strong leadership... ...market. As deep learning models become more versatile...Local areaWork from homeRelocationFlexible hours$253.1k - $342.3k
..., you’ll support the development and management of Compute, Database,... ...team that builds AWS Neuron, the software stack that runs all the leading AI models on the AWS Inferentia... ...Manager in the AWS Neuron SDK organization you will... ...and libraries that enable users to develop and...Local areaFlexible hours- ...the framework-layer inference product strategy for... ...help shape how AMD’s software ecosystem enables efficient, reliable,... ...competitive large-scale model deployment. You will... ...orchestration, memory management, and performance while... ...systems such as llm-d, along with memory...Remote workShift work
$192k - $278k
...product strategy for AI model serving safety... ...for AI model inference, including real-time... ...grade safety controls enabling custom threshold... ...abuse triage and manage regulatory escalation... ...Language Model (LLM) interfaces into... ...relationships, commercial co-development contracts, and...- ...seeking a Senior Product Manager to drive strategy and... ...AMD’s open-source GPU software stack, with a specific focus on large-scale model inference on AMD Instinct™ and... ...and the frameworks that enable it. You will influence... ...understanding of LLM inference, including attention...Remote work
- ...team in: research, design, development, and deployment of... ...integrate large language models (LLMs) and other state-of... ...proven by a track record of software artifacts or academic... ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in...Full timeWork experience placement
$117.7k - $221.4k
...and robotics depends not only on stronger models, but also on better infrastructure for... ...the data processing, featurization, and inference foundations that power scalable world understanding... ...{or other frequency dictated by your manager}. Relocation...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$192.2k - $260k
...'s Delivery Foundation Model team, where you'll work... ...foundation models that enable delivery of billions of... ...amounts of Amazon data and infer at Amazon scale, taking... ...of foundation model development, from multimodal... ...judgment, effectively manage stress and work safely...Local areaWorldwideFlexible hours$136.8k - $277.2k
...company\'s risk management, governance and... ...products that enable and empower continuous... ...Model Evaluation & Audit... ...Professional Development: Continue to develop... ...transformer-based LLM architectures (... ...pipelines, and inference behaviors.... ...end and back end software development skills...Temporary workLocal area$152k - $241.5k
...computing platforms, software, and AI systems that enable autonomous vehicles... ...to accelerate the development, integration,... ...end-to-end driving models across large-scale... ...output, timing, state-management, and runtime interface... ...platforms.Optimize model inference and surrounding...Full time- ...group is to create cutting‑edge technology enabling next‑generation physical AI, with... ...of flexibility and trust our employees to manage their schedules responsibly. This may include... ...research on pretraining world-action foundation model with various world modalities including...For contractorsFor subcontractorCasual workInternshipWork at officeRemote workDay shift
$165.45k - $259.75k
Principal Software Architect Description... ...workflows, AI-enabled capabilities,... ...language models, generative AI... ...Edge AI, local inference, and model optimization... ..., UX, product management, hardware... ...experiences, LLM and agentic... ...backend and systems development experience...Full timeTemporary workLocal areaRelocationFlexible hoursShift work$171.6k - $222.2k
...learning and large language models.We leverage advanced... ....We are pioneering the development of robotics foundation models that: - Enable unprecedented... ...Experience in professional software developmentAmazon is an... ...judgment, effectively manage stress and work safely...Local areaWorldwideFlexible hours$272k - $431.25k
...generation of interactive world-model systems. With this release,... ...for fidelity, real-time inference performance in world models.... ...Success means more than shipping software: it means high-quality solutions... ..., CI/CD, and release management.Partner with research, simulation...Full time- ...Institute of Foundation Models We are a dedicated... ...understanding, using, and risk-managing foundation models. Our... ...challenges in AI development. You will participate... ...model modularity, and inference optimization. Build... ..., and open-source software. Represent MBZUAI at...
- ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-... ...computation. Cerebras works with the leading model labs, global enterprises, and cutting-... ...People who are serious about software make their own hardware. At Cerebras, we...Full time
$224k - $356.5k
...work directly with software solution providers... ...Omniverse and AI enabled-solutions.Develop... ...innovation and the development of next-generation... ...on AV End-to-End models and GenAI model development... ...in deploying LLM models at scale on... ...and optimize inference latency and throughput...Full time$165k - $185k
...Valley focuses on Foundation Models, Big Data Visual Analytics,... ...Foundation Model Powered AI Enablers group plays a pivotal role in... ...expert insights to the management team in relevant technology... ...researchHands-on experience in product development in the above-mentioned areas...Work experience placementWorldwide- ...to spearhead the development of next-generation... ...autonomous and edge-enabled systems. The CTO... ...deployment of AI/ML or LLM-enabled systems on... ...large language model technologies, and... ...will have a strong software engineering... ...with embedded AI inference frameworks and edge...
$272k - $431.25k
...the future, we’re generating it! Our world model team is pushing the boundaries of... ...AI. We are looking for a Senior Research Manager to lead world-model evaluation and benchmarking... ...evaluations, where access to model internals enables deeper diagnostics, causal analysis, and...Full time- ...benefits of AI every day—enabling medical research,... ...- a cutting-edge 'AI Software Solutions Team'. Specialized... ...tuning large language models to unlock... ...and Python "strong development experience is at least... ...major DL framework in inference, fine tuning and/or training...
$219k - $351k
...bandwidth business. As models scale past what any... ...the core product of AI inference, not an afterthought.\... ...Samsung Cognos, AI memory software that moves model state... ...Dynamo, TensorRT-LLM, llama.cpp-class engines... ...including their memory-management internals, not just...Work at officeRemote workFlexible hours$140k - $164.75k
...driving the transformation to AI-enabled software-defined vehicles.... ...software integration and prototype development, you will havethe... ...Integrate software modules and SDK at the system or application... ...technology advantages. Customer Management: Provide integration support...Work experience placementWork at officeImmediate startRemote workWorldwideFlexible hoursShift work$245k - $325k
Software Architect San Jose, California, United States... ..., from chip to model, optimized for enterprise... .... Our SambaStack inference serving platform is... ...serving platform enables seamless deployment, management, and scaling of foundation... ...challenges of LLM workloads at production...Full timeTemporary workLocal areaFlexible hours- ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-... ...computation. Cerebras works with the leading model labs, global enterprises, and cutting-... ...decisions, mentor engineers, and write software and automation alongside the team....Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Development Manager, LLM Inference Model Enablement, Neuron SDK. Be the first to apply!
- embedded software Cupertino, CA
- entry level software sales Cupertino, CA
- software technology Cupertino, CA
- software implementation project manager Cupertino, CA
- software support Cupertino, CA
- software engineer - cloud services Cupertino, CA
- bank software Cupertino, CA
- ultimate software Cupertino, CA
- remote software sales Cupertino, CA
- healthcare software sales Cupertino, CA



