Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron

$193.3k - $261.5k

Amazon Locker

We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.As a Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Trainium accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation.Key job responsibilities* Deliver high-performance models using distributed inference libraries* Drive technical excellence in performance optimization and system reliability across the Neuron ecosystem* Mentor team members and provide technical leadership across multiple work streams* Drive architectural decisions that impact the entire Neuron serving stack* Collaborate with customers, product owners, and engineering teams to define technical strategy* Author technical documentation, design proposals, and architectural guidelinesA day in the lifeYou'll lead critical technical initiatives while mentoring team members. You'll collaborate with cross-functional teams of applied scientists, system engineers, and product managers to architect and deliver state-of-the-art inference capabilities. Your day might involve:* Leading design reviews and architectural discussions* Debugging complex performance issues across the stack in collaboration with the compiler and runtime teams* Mentoring junior engineers on system design and model optimization across model enablement teams* Driving technical decisions that shape the future of Neuron's inference stackAbout the teamThe inference model enablement team releases its models in the vLLM Neuron plugin: qualifications- 5+ years of programming using a modern programming language such as Java, C++, or C#, including object-oriented design experience- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience- 5+ years of non-internship professional software development experience- Experience as a mentor, tech lead or leading an engineering teamPreferred qualification - Master's degree in computer science or equivalent- Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model executionAmazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 193,300.00 - 261,500.00 USD annually

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron in Cupertino, CA vacancy
  • $193.3k - $261.5k

     ...Web Services (AWS) builds AWS Neuron, the software development kit used to...  ...and Trainium ML accelerators....  ...PyTorch and JAX enabling unparalleled ML inference and training...  ...range of models and supporting...  ...boundary, our engineers build systematic...  ...possible in AI acceleration.... 
    Amazon Web Service
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $212.7k - $287.7k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud...  ....As an SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize...  ...* Report on status of development, quality, operations, and... 
    Amazon Web Service
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $165.2k - $223.6k

    AWS Neuron is the complete software stack for the AWS Inferentia and...  ...the Software Development Engineer for the Neuron...  ...applications and AI accelerators....  ...performance of ML Kernels and ML...  ...distributed training and inference solutions. This...  ...and enable them to take on... 
    Amazon Web Service
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    6 hours ago
  • $184k - $287.5k

     ...and motivated software engineers to join us and build AI inference systems that serve...  ...large-scale models with extreme efficiency...  ...the field of ML Systems; survey...  ...platforms (AWS/GCP/Azure),...  ...AI research and development to create groundbreaking...  ...that enable anyone to harness... 
    Amazon Web Service
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $165.2k - $223.6k

     ...acquired by AWS in 2015 and is...  ...integrated. AWS Neuron isthe complete software stack for the...  ...and Trainium ML accelerators...  ...large-scale AI workloads on...  ...for a Software Development Engineer to build...  ...integrations that enable customers to...  ...training and inference workloads... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    6 hours ago
  • $193.3k - $261.5k

     ...designs silicon and software that accelerates...  ...software stacks enable us to tackle...  ...change the world. AWS Neuron is the complete...  ...Senior Software Engineer to join our ML Distributed...  ...responsible for the development, enablement, and...  ...large scale ML model training across... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $117.7k - $221.4k

     ...and cost efficient for embodied AI systems. We believe the next...  ...depends not only on stronger models, but also on better infrastructure...  ...model reflects how Cola engineers think: build durable intermediate...  ...processing, featurization, and inference foundations that power... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $250k - $344.5k

     ...Security (NetSec) Engineering – Our team is at the...  ...customer-facing AI-enabled solutions designed...  ...specialized group of software engineers...  ...you will lead the development of a diverse portfolio...  ...that leverage AI/ML to solve real-world...  ...native infrastructure (AWS/GCP, Kubernetes,... 
    Amazon Web Service

    Palo Alto Networks, Inc.

    Santa Clara, CA
    3 days ago
  • $193.3k - $261.5k

     ...an integral part of AWS and develops hardware and software components that are...  ...experience.The AWS Neuron Collectives team is...  ...seeking a Software Engineer to optimize...  ...powering the frontier AI models being trained today...  ...rounded professional and enable them to take on... 
    Amazon Web Service
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $165.2k - $223.6k

    As a Neuron Collectives Software Developer, you will...  ...operations to scale AI compute...  ...driver development* Work closely...  ...crucial part of AWS, is...  ...Machine Learning (ML) and High-Performance...  ..., hardware engineers, RTL...  ...stack that enables collective operations...  ...frontier models that power... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $127.1k - $185k

     ...talented early-career engineer to join our team...  ...for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive...  ...running on custom AWS hardware - Build and...  ...Familiarity with Linux development environments and... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $50k - $120k

     ...Altimate AI, founded in 2...  ...multiple language models and a custom-...  ...graph, we enable contextually...  ...powered data engineering revolution. You...  ...of hands-on ML/AI experience...  ...architectures API development expertise (...  ...(AWS, Kubernetes)...  ...scale training, inference, and multi-agent... 
    Amazon Web Service
    Full time
    Worldwide

    Pa Early Stage Partners

    Sunnyvale, CA
    10 hours ago
  • $193.3k - $261.5k

     ...tools for the Neuron ML accelerators...  ...and software teams to ensure...  ...performance engineers to develop and...  ...including training, inference and runtime....  ...Experiences AWS values...  ...support the development and...  ...generative AI services and...  ...s all being enabled by AWS Neuron... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $165.2k - $223.6k

     ...part of the AWS Specialist and...  ...recruiting, development, and growth of...  ...challenging engineering domains, and...  ...the future of AI-driven cloud...  ...are seeking a Software Development Engineer...  ...security, AI/ML, distributed...  ...that enable rapid prototyping...  ...language models, autonomous agents... 
    Amazon Web Service
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  •  ...potential of generative AI to power the...  ...the forefront of software and hardware...  ...System Software Engineer, AI Inference ExecutionWhat you...  ...responsible for the development, enhancement, and...  ...other software (ML and compilers) and...  ...inference servers/model serving frameworks... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $229.9k - $262.4k

     ...Overview Sr. Lead AI Engineer (Inference Optimization, FM...  ...applications of AI & ML are bringing...  ...customers.  Our AI models and platforms empower...  ..., and support AI software components including...  ...such as AWS Ultraclusters, Huggingface...  ...software, and AI enable you to see and exploit... 
    Amazon Web Service
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    more than 2 months ago
  • $120.75k - $161k

     ...generation of Agentic AI products that help millions...  .... We’re looking for a Software Engineer to design and build...  ..., designers, and AI/ML engineers to bring intelligent...  ...and APIs that support model inference, orchestration, and...  ...with cloud platforms (AWS, GCP, or Azure)... 
    Amazon Web Service
    Work at office
    Remote work
    Flexible hours

    Eightfold

    Santa Clara, CA
    1 day ago
  • $165.2k - $223.6k

     ...and Agentic AI ?The team manages...  ...), and the AWS Glue Python...  ...that enables customers to...  ...2 months of engineering effort for customers...  ...managed Model Context Protocol...  ...-quality software applying...  ...professional software development experience-...  ..., training/inference lifecycles,... 
    Amazon Web Service
    Internship
    Local area
    Remote work
    Flexible hours

    Amazon

    East Palo Alto, CA
    1 day ago
  •  ...RoboForce is an AI robotics company...  ...looking for a Senior Software Engineer to build...  ...infrastructure that enables large-scale model training, validation...  ...into on-robot inference stacks. Requirements...  ...C++, Python, and ML frameworks (e.g.,...  ...provider (GCP, AWS, Azure) and... 
    Amazon Web Service
    Full time
    Work at office
    Visa sponsorship

    RoboForce

    Milpitas, CA
    10 hours ago
  • $152k - $241.5k

     ...are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at...  ...researchers to make their ML models more efficient...  ...usage of hardware, software, and infrastructure...  ..., training & inference performance end to...  ...computing platforms (e.g., AWS, GCP, Azure) in... 
    Amazon Web Service
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...distributed systems (pre-AI experience...  ...tools, or MCP (Model Context...  ...Proficient for backend development...  ...experience with AWS/GCP/Azure - cost...  ...Hands-On Engineer Not just an...  ...Streaming inference and async agent...  ...compression ML observability tools... 
    Amazon Web Service

    ClifyX

    Sunnyvale, CA
    1 day ago
  • $151.8k - $332.2k

     ...expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will...  ..., C/C++; familiarity with ML frameworks such as PyTorch and...  ...and linguistic representations, enabling unified modeling for speech understanding... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    4 days ago
  • $144.25k - $256.25k

     ...benefitsJob Function: Engineering &...  ...organization enables and accelerates...  ...American Express, AI is reshaping the...  ...on agentic AI development: designing...  ...infrastructure, inference, and model gatewaysEvaluation...  ...: AWS and/or GCP, KubernetesDistributed...  ...or advanced ML... 
    Amazon Web Service
    Work at office
    Visa sponsorship
    3 days per week

    American Express

    Palo Alto, CA
    6 hours ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...platforms. As a Senior AI/ML Research Engineer, you...  ...fine-tune the foundation models—VFMs, VLMs, and VLA models...  ...multimodal models—that enable the system to perceive...  ...ML research, robotics, software, and data engineering to... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    2 days ago
  • $105k - $115k

     ...world-class end-to-end engineering solutions by...  ...dimensional approach enables us to solve the most...  ...with by utilizing Gen AI or other machine...  ...platformsAI Engineer/ ML Engineer with...  ...some knowledge in aws or some cloud.Assess...  ...6-07-31Profession: Software & DigitalEmployment... 
    Amazon Web Service
    Temporary work

    Quest Global Services

    Sunnyvale, CA
    1 day ago
  •  ...from fraud to enabling companies to...  ...can thrive.AI Engineer — Customer...  ...operate the core ML/AI systems...  ..., own model lifecycle and...  ...APIs (scalable inference, caching, batching...  ...Full-Stack Development: Design,...  ...on Azure or AWS, leveraging...  .... 10+ years software engineering... 
    Amazon Web Service
    Full time
    Local area

    F5 Networks

    San Jose, CA
    4 days ago
  • $170.6k - $261.3k

     ...team: The AV ML Infra team at...  ...unique demands of AI and ML...  ...and more. We enable scalable and...  ...productivity of ML engineers, and drive...  ...Validation & Inference: Ensures robust model performance by...  ...end-to-end software products, owning...  ...: Full-Stack Development: Design,... 
    Amazon Web Service
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $227.5k - $300k

     ...transformation to AI-enabled software-defined vehicles....  ...a Senior Staff AI Engineer with a combination...  ...rigor to lead the development of an Agentic Framework...  ...allowing for model-agnostic routing and...  ..., traditional ML models, etc. (AI depth...  ...platforms (e.g., AWS, Azure, Google Cloud... 
    Amazon Web Service
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    4 days ago
  • $149.52k - $175.9k

     ...financial decisions and enabling the communities we...  ...of enterprise AI and Generative AI...  ....Partner with engineering, product, architecture...  ...in Large Language Models (LLMs), Agentic AI...  ...supporting AI/ML Platforms, MLOps,...  ...OpenAI), and modern software engineering practices... 
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Cupertino, CA
    2 days ago
  • $114.1k - $214.95k

     ...Document Cloud’s AI team is building the...  ...’re looking for a Software Development Engineer to help build and...  ...and pipelines that enable our Machine...  ...features and the ML pipelines that power...  ...data pipelines for model evaluation, prompt...  ...cloud platforms (AWS, GCP, or Azure) and... 
    Amazon Web Service
    Full time
    Contract work
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron. Be the first to apply!