Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Development Manager, LLM Inference Model Enablement, Neuron SDK

$212.7k - $287.7k

Amazon Locker

We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scalemachine learning accelerators. Join us to optimize LLMs to run really fast on the Trainium hardware.As an SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.The ideal candidate will have a strong background in LLM model architectures, model performance optimizations, and inference techniques, such as delivering high-performance models using distributed inference libraries. You should be capable of managing demanding, fast-changing priorities. You should have a strong technical ability to understand and deliver as part of a vertically integrated system stack consisting of the PyTorch inference library, Neuron compiler, runtime, and collectives.Key job responsibilities* Management and execution against project plans and delivery commitments* Manage the day-to-day activities of the engineering team* Management of resources, staffing, mentoring, and maintaining a best-of-class engineering team* Report on status of development, quality, operations, and model performance to managementA day in the lifeYou will work with your senior management and technical leaders to define the model enablement and performance optimization for the latest SOTA LLMs, build and deliver them to customers.Meanwhile, lead the team to continue improving the model onboarding experience, as well as enhancing inference usability and quality for Neuron-supported models.You will manage changing priorities as new models and new technologies emerge, and you adapt your team’s work to manage them. You will dive deep to help your team solve technical challenges.Basic qualifications- 3+ years of engineering team management experience- 7+ years of working directly within engineering teams experience- 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience- Experience partnering with product or program management teamsPreferred qualification - Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy- Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineersAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 212,700.00 - 287,700.00 USD annually

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Software Development Manager, LLM Inference Model Enablement, Neuron SDK in Cupertino, CA vacancy
  • $165.2k - $223.6k

     ...AWS) builds AWS Neuron, the software development kit used to...  ...The AWS Neuron SDK, developed by...  ...PyTorch and JAX enabling unparalleled ML inference and training...  ...wide range of models and supporting...  ...wide variety of LLM model families...  ..., and product managers to deliver... 
    Suggested
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    29 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon'...  ...optimize the latest models to run really fast on...  ...hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will...  ...engineers, and product managers to architect and deliver... 
    Suggested
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $212.7k - $287.7k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...of Software Development for the...  ...engineers and managers to help design...  ...of latest ML models at scale on...  ...engineers focused on enabling new ML...  ...on the Neuron SDK / Trainium...  ...training and inference solutions. This... 
    Suggested
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  •  ...leading training and inference speeds; over 10...  ...the leading model labs, global...  ...Model Scaling team enables state-of-the-...  ...as well as development of high-performance...  ..., product management, and AI research...  ...support for emerging LLM architectures...  ...hardware/software co-design through... 
    Suggested

    Cerebras Systems

    Sunnyvale, CA
    7 days ago
  • $212.7k - $287.7k

     ...for ML in the cloud. This is all enabled by the AWS Neuron Software Development Kit (SDK), which includes an ML compiler,...  ...a talented Software Development Manager with strong leadership and mentoring...  ...to market. As deep learning models become more versatile, using compiler... 
    Suggested
    Local area
    Work from home
    Relocation
    Flexible hours
    Day shift

    Amazon

    Cupertino, CA
    a month ago
  • $212.7k - $287.7k

     ...delivers best-in-class ML inference performance at the...  ...cloud. This is all enabled by edge software stack, the AWS Neuron Software Development Kit (SDK), which includes an ML...  ...SW Engineering Manager with strong leadership...  ...market. As deep learning models become more versatile... 
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $253.1k - $342.3k

     ..., you’ll support the development and management of Compute, Database,...  ...team that builds AWS Neuron, the software stack that runs all the leading AI models on the AWS Inferentia...  ...Manager in the AWS Neuron SDK organization you will...  ...and libraries that enable users to develop and... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...the framework-layer inference product strategy for...  ...help shape how AMD’s software ecosystem enables efficient, reliable,...  ...competitive large-scale model deployment. You will...  ...orchestration, memory management, and performance while...  ...systems such as llm-d, along with memory... 
    Remote work
    Shift work

    AMD

    Santa Clara, CA
    a month ago
  • $192k - $278k

     ...product strategy for AI model serving safety...  ...for AI model inference, including real-time...  ...grade safety controls enabling custom threshold...  ...abuse triage and manage regulatory escalation...  ...Language Model (LLM) interfaces into...  ...relationships, commercial co-development contracts, and... 

    Google

    Sunnyvale, CA
    4 days ago
  •  ...seeking a Senior Product Manager to drive strategy and...  ...AMD’s open-source GPU software stack, with a specific focus on large-scale model inference on AMD Instinct™ and...  ...and the frameworks that enable it. You will influence...  ...understanding of LLM inference, including attention... 
    Remote work

    AMD

    Santa Clara, CA
    1 day ago
  •  ...team in: research, design, development, and deployment of...  ...integrate large language models (LLMs) and other state-of...  ...proven by a track record of software artifacts or academic...  ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM). Research experience in... 
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    1 day ago
  • $117.7k - $221.4k

     ...and robotics depends not only on stronger models, but also on better infrastructure for...  ...the data processing, featurization, and inference foundations that power scalable world understanding...  ...{or other frequency dictated by your manager}. Relocation... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $192.2k - $260k

     ...'s Delivery Foundation Model team, where you'll work...  ...foundation models that enable delivery of billions of...  ...amounts of Amazon data and infer at Amazon scale, taking...  ...of foundation model development, from multimodal...  ...judgment, effectively manage stress and work safely... 
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    11 days ago
  • $136.8k - $277.2k

     ...company\'s risk management, governance and...  ...products that enable and empower continuous...  ...Model Evaluation & Audit...  ...Professional Development: Continue to develop...  ...transformer-based LLM architectures (...  ...pipelines, and inference behaviors....  ...end and back end software development skills... 
    Temporary work
    Local area

    Jobleads-US

    San Jose, CA
    1 day ago
  • $152k - $241.5k

     ...computing platforms, software, and AI systems that enable autonomous vehicles...  ...to accelerate the development, integration,...  ...end-to-end driving models across large-scale...  ...output, timing, state-management, and runtime interface...  ...platforms.Optimize model inference and surrounding... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...group is to create cutting‑edge technology enabling next‑generation physical AI, with...  ...of flexibility and trust our employees to manage their schedules responsibly. This may include...  ...research on pretraining world-action foundation model with various world modalities including... 
    For contractors
    For subcontractor
    Casual work
    Internship
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    6 days ago
  • $165.45k - $259.75k

    Principal Software Architect Description...  ...workflows, AI-enabled capabilities,...  ...language models, generative AI...  ...Edge AI, local inference, and model optimization...  ..., UX, product management, hardware...  ...experiences, LLM and agentic...  ...backend and systems development experience... 
    Full time
    Temporary work
    Local area
    Relocation
    Flexible hours
    Shift work

    HP

    Palo Alto, CA
    4 days ago
  • $171.6k - $222.2k

     ...learning and large language models.We leverage advanced...  ....We are pioneering the development of robotics foundation models that: - Enable unprecedented...  ...Experience in professional software developmentAmazon is an...  ...judgment, effectively manage stress and work safely... 
    Local area
    Worldwide
    Flexible hours

    Amazon

    Sunnyvale, CA
    4 days ago
  • $272k - $431.25k

     ...generation of interactive world-model systems. With this release,...  ...for fidelity, real-time inference performance in world models....  ...Success means more than shipping software: it means high-quality solutions...  ..., CI/CD, and release management.Partner with research, simulation... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  •  ...Institute of Foundation Models We are a dedicated...  ...understanding, using, and risk-managing foundation models. Our...  ...challenges in AI development. You will participate...  ...model modularity, and inference optimization. Build...  ..., and open-source software. Represent MBZUAI at... 

    Institute of Foundation Models

    Sunnyvale, CA
    24 days ago
  •  ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-...  ...computation. Cerebras works with the leading model labs, global enterprises, and cutting-...  ...People who are serious about software make their own hardware. At Cerebras, we... 
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $224k - $356.5k

     ...work directly with software solution providers...  ...Omniverse and AI enabled-solutions.Develop...  ...innovation and the development of next-generation...  ...on AV End-to-End models and GenAI model development...  ...in deploying LLM models at scale on...  ...and optimize inference latency and throughput... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $165k - $185k

     ...Valley focuses on Foundation Models, Big Data Visual Analytics,...  ...Foundation Model Powered AI Enablers group plays a pivotal role in...  ...expert insights to the management team in relevant technology...  ...researchHands-on experience in product development in the above-mentioned areas... 
    Work experience placement
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    a month ago
  •  ...to spearhead the development of next-generation...  ...autonomous and edge-enabled systems. The CTO...  ...deployment of AI/ML or LLM-enabled systems on...  ...large language model technologies, and...  ...will have a strong software engineering...  ...with embedded AI inference frameworks and edge... 

    Confidential

    San Jose, CA
    2 days ago
  • $272k - $431.25k

     ...the future, we’re generating it! Our world model team is pushing the boundaries of...  ...AI. We are looking for a Senior Research Manager to lead world-model evaluation and benchmarking...  ...evaluations, where access to model internals enables deeper diagnostics, causal analysis, and... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  •  ...benefits of AI every day—enabling medical research,...  ...- a cutting-edge 'AI Software Solutions Team'. Specialized...  ...tuning large language models to unlock...  ...and Python   "strong development experience is at least...  ...major DL framework in inference, fine tuning and/or training... 

    AMD

    Santa Clara, CA
    7 days ago
  • $219k - $351k

     ...bandwidth business. As models scale past what any...  ...the core product of AI inference, not an afterthought.\...  ...Samsung Cognos, AI memory software that moves model state...  ...Dynamo, TensorRT-LLM, llama.cpp-class engines...  ...including their memory-management internals, not just... 
    Work at office
    Remote work
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    5 days ago
  • $140k - $164.75k

     ...driving the transformation to AI-enabled software-defined vehicles....  ...software integration and prototype development, you will havethe...  ...Integrate software modules and SDK at the system or application...  ...technology advantages. Customer Management: Provide integration support... 
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    2 days ago
  • $245k - $325k

    Software Architect San Jose, California, United States...  ..., from chip to model, optimized for enterprise...  .... Our SambaStack inference serving platform is...  ...serving platform enables seamless deployment, management, and scaling of foundation...  ...challenges of LLM workloads at production... 
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova Systems

    San Jose, CA
    5 days ago
  •  ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-...  ...computation. Cerebras works with the leading model labs, global enterprises, and cutting-...  ...decisions, mentor engineers, and write software and automation alongside the team.... 
    Full time

    Cerebras

    Sunnyvale, CA
    11 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Development Manager, LLM Inference Model Enablement, Neuron SDK. Be the first to apply!