Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Software Development Engineer - AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs

$193.3k - $261.5k

Amazon Locker

Every token a large language model generates depends on data reaching the right accelerator at the right moment. As AI models outgrow any single chip, the network between accelerators becomes the bottleneck that decides how fast — and how affordably — the world's largest models can serve real users. That network layer is what our team builds.We're looking for an engineer to work at the frontier of disaggregated inference: splitting LLM serving into separate prefill and decode pools and moving the model's KV cache between them at the absolute limit of what the hardware allows. Get it right and users get answers in milliseconds; get it wrong and the fastest accelerators in the world sit idle waiting on data. You'll own pieces of the high-speed transfer path that make the difference, and you'll measure your success in how close you run to the theoretical peak of the machine.In this role you will:Build and optimize the low-level data-movement software across accelerators, servers, and heterogeneous memory — over AWS's highest-performance network fabric.Push performance to the hardware roofline: profile, find the real bottleneck, and close the gap between "it works" and "it runs as fast as physics permits."Work across the stack — from kernel and network transport up to the inference frameworks — and partner with teams building the chips, runtime, and models.Deliver features that run on our largest clusters, for our largest customers, serving the largest AI models in production.What we're looking for:Strong C/C++ and a love for low-level, performance-critical systems — solid command of Linux, kernels, memory, and writing fast code.The instinct to ask "how fast could this possibly go?" and the rigor to measure it.Experience with high-speed networking or HPC interconnects (RDMA, InfiniBand, libfabric, UCX, NIXL, MPI) is valued highly; embedded-systems experience is a plus.Prior AI/ML experience is welcome but not required — if you're a great systems engineer, we'll teach you the ML side.If you like solving genuinely hard problems, working shoulder-to-shoulder with HPC and ML customers, iterating fast, and shipping at a scale few places can offer, come join us. This is a role on the leading edge of AI/ML infrastructure.About the team: You'd be joining Annapurna Labs, an integral part of AWS. Annapurna designs the hardware and software building blocks behind EC2 — every EC2 instance runs on hardware we designed. We specialize in the chips, systems, and software that optimize the AWS customer experience, and this team sits where cutting-edge AI meets the silicon and the network underneath it.A day in the lifeAnnapurna Labs, a crucial part of AWS, is responsible for developing hardware and software components for EC2 infrastructure. Our team focuses on building networking solutions that for Machine Learning (ML) and High-Performance Computing (HPC) workloads on AWS.We have mixed discipline orgs, you’d be working side by side with infrastructure experts, hardware engineers, RTL engineers, scientists & architects. Our workforce spans the globe and is truly international, you’ll find yourself working side by side with individuals from numerous countries. We take mentorship seriously, you can both expect senior mentorship and will be expected to mentor new and junior engineers. The pace is fast as we work on the latest advancements of AI/ML, but we take the time to bond as a team and enjoy the successes. We offer flexibility in working hours, and respect WLB as a core org tenet. The team enjoys working with numerous principal-level engineers and closely with directors, career growth opportunities are certainly available. This is a role where you will always be encouraged to keep learning, the AI/ML field is fast moving and constantly evolving.Basic qualifications- 5+ years of non-internship professional software development experience- 5+ years of programming with at least one software programming language experience- 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience- Experience as a mentor, tech lead or leading an engineering team- Must have C/C++ Coding ExperiencePreferred qualification - Bachelor's degree in computer science or equivalent- ML Communications (NCCL, NIXL, NVSHMEM)Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Cupertino - 193,300.00 - 261,500.00 USD annually

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Sr. Software Development Engineer - AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs in Cupertino, CA vacancy
  • $193.3k - $261.5k

     ...seeking an experienced engineer to work on distributed AI/ML systems. This role...  ...with high-speed networking or HPC...  ...would be joining is Annapurna Labs, an integral part...  ...develops hardware and software components that are...  ...professional software development experience- 5+... 
    Senior
    Network
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $165.2k - $223.6k

     ...seeking an experienced engineer to work on distributed AI/ML systems. This role...  ...with high-speed networking or HPC...  ...would be joining is Annapurna Labs, an integral part...  ...develops hardware and software components that are...  ...professional software development experience- 2+... 
    Network
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $206.9k - $279.9k

     ...custom kernel development and...  ...acceleration software. AWS Neuron...  ...best-in-class ML performance...  ...and influence engineering discussions...  ...in-class ML inference performance...  ...About Amazon Annapurna Labs:Amazon Annapurna...  ...ML chips, in networking and security...  ...generative AI services and... 
    Network
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $208.3k - $281.8k

     ...training and inference of frontier...  ...Neuron is the software stack for...  ...generative AI workloads with...  ...training AI/ML ecosystem and...  ...with engineering teams building...  ...About Amazon Annapurna Labs Amazon Annapurna...  ...chips, in networking and security...  ...support the development and management... 
    Network
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $176.6k - $239k

     ...Annapurna Labs was a startup company acquired...  ...silicon engineering, hardware...  ...verification, software, and...  ...and Trainium ML Accelerators...  ...in-class ML inference performance...  ...Neuron Software Development Kit (SDK), which...  ...in networking and security...  ...and now in AI and Machine... 
    Senior
    Network
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $193.3k - $261.5k

     ...Generative AI on AWS. The...  ...best-in-class ML inference performance...  ...cutting edge software stack, the AWS...  ...Software Development Kit (SDK), which...  ...the Amazon Annapurna Labs team is...  ...including silicon engineering, hardware...  ...takes neural network descriptions...  ....You: As a Sr. Machine Learning... 
    Senior
    Network
    Internship
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    14 days ago
  • $127.1k - $185k

     ...looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI...  ...scalability)- Familiarity with Linux development environments and toolchainsPreferred... 
    Network
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    22 days ago
  • $165.2k - $223.6k

    The Annapurna Labs team at Amazon Web Services (AWS...  ...AWS Neuron, the software development kit used to...  ...Inferentia and Trainium ML accelerators....  ...unparalleled ML inference and training performance...  ...boundary, our engineers build systematic...  ...what's possible in AI acceleration.As... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    29 days ago
  • $165.2k - $223.6k

    Annapurna Labs is an integral part of AWS and develops hardware and software components that are critical building blocks...  ...seeking a Software Engineer to optimize...  ...powering the frontier AI models being trained...  ...full software/hardware/networks development life cycle, including... 
    Network
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $127.1k - $185k

    Annapurna Labs was a startup acquired by AWS in 201...  ...org spans silicon engineering, hardware design and verification, software, and operations. We...  ...and Trainium ML Accelerators, and...  ...looking for a Software Development Engineer to help build...  ...on custom AI accelerators. You'... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $183k - $247.6k

    Annapurna Labs designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions...  ...custom designed machine learning inference datacenter server. Our success...  ...our team members develop your engineering expertise so you feel empowered to... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $193.3k - $261.5k

     ...seeking an experienced engineer and technical...  ...team that owns the network stack for EC2 distributed AI/ML systems. The team...  ...would be joining is Annapurna Labs, an integral part...  ...hardware and software components that are...  ...of full software development life cycle, including... 
    Network
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $193.3k - $261.5k

     ...learning training and inference clusters. Our...  ...the low-level software stack that...  ...together across a network.We're looking...  ...Systems Software Engineer who wants to...  ...system software development, and enable architectural...  ...As part of the ML accelerator...  ...Labs, our organization... 
    Senior
    Network
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $193.3k - $261.5k

    The Annapurna Labs team at Amazon Web Services (AWS...  ...AWS Neuron, the software development kit used to...  ...for AWS's custom ML accelerators. Working...  ...software boundary, our engineers craft high-...  ...what's possible in AI acceleration.The...  ...unparalleled ML inference and training performance... 
    Senior
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $193.3k - $261.5k

    Annapurna Labs designs silicon and software that accelerates innovation. Our custom chips, accelerators, and...  ...are seeking a Senior Software Engineer to join our ML Distributed Training team.In...  ...you will be responsible for the development, enablement, and performance optimization... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    19 days ago
  • $174k - $252k

     ...technical design, development, and optimization of software components...  ...Language Model (LLM) inference serving on GDC. This...  ...techniques like disaggregated serving, speculative...  ...Kubernetes Engine (GKE)), networking infrastructure, and...  ...development, including AI/ML applications.... 
    Senior
    Network

    Google

    Sunnyvale, CA
    3 days ago
  • $208.3k - $281.8k

     ...training and inference of frontier...  ...Neuron is the software stack for...  ...generative AI workloads with...  ...optimize, and tune ML models for...  ...with engineering teams building...  ...Marketing, Business Development, and...  ...About Amazon Annapurna Labs Amazon...  ...ML chips, in networking and security... 
    Network
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    7 days ago
  • $92k - $135k

     ...Essential Cloud for AI™. Built for...  ...Trusted by leading AI labs, startups, and global...  ...You'll Do: Join the Inference team to ship production...  ...from experienced engineers. About the role:...  ..., algorithms, and networked services....  ...a microservice or ML inference demo.... 
    Network
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    18 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators...  ...to run really fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  •  ...world's largest AI chip, 56...  ...training and inference speeds; over...  ...leading model labs, global enterprises...  ...a Staff Engineer to help lead,...  ...maintain production software, with...  ..., continuous development, observability, security, networking, debugging, and...  ...Partner with ML, Product, Infrastructure... 
    Senior
    Network

    Cerebras Systems

    Sunnyvale, CA
    21 days ago
  • $184k - $287.5k

     ...skilled and motivated software engineers to join us and build AI inference systems that serve large...  ...parallelism, prefill-decode disaggregation.Develop, optimize, and...  ...for the field of ML Systems; survey recent...  ...advance AI research and development to create groundbreaking... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  •  ...computing experiences—from AI and data centers, to...  ...:We are hiring AI / ML Platform Engineers to build the platform...  ...training and inference, experiment tracking,...  ...infrastructure, and hardware/software tooling, and you can...  ..., storage systems, networking, containers, and... 
    Senior
    Network

    AMD

    Santa Clara, CA
    a month ago
  • $124.5k - $272k

     ...in modern security and networking for the cloud and AI era. We secure and accelerate...  ..., its Zero Trust Engine, and the powerful NewEdge...  ...Scientist, you own the inference and optimization layer that...  ...with 4+ years hands-on in ML/AI (model development, fine-tuning, and... 
    Senior
    Network

    Netskope

    Santa Clara, CA
    21 days ago
  • $183k - $247.6k

     ...that are used to power today’s AI workloads in datacenters all around the world. As a Sr. SoC Power Engineer, you’ll contribute to the...  ...designers, verification engineers, software teams, and Physical Design...  ...power measurements in the lab and correlate back to simulations... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $183k - $247.6k

     ...used to power today’s AI workloads in datacenters...  ...experienced Design Verification Engineers to build the next...  ...- 8+ YOE in testbench development including: stimulus,...  ...bench, FPGA, emulator, software environments, and system...  ...verifying complex CPU, GPU, or ML accelerator... 
    Senior
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $184k - $287.5k

     ...capability gains in AI today. It is...  .... RL requires inference, rollout...  ...RL Frameworks engineering team to develop...  ...spans the full software stack, from collaborating...  ...and labs pushing the frontier...  ...with NVIDIA's networking, math library,...  ...infrastructure, or ML systems engineeringStrong... 
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    23 days ago
  • $183k - $247.6k

     ...organization, you’ll support the development and management of...  ...their cloud services.Annapurna Labs (our organization...  ...) designs silicon and software that accelerates innovation...  ...a Hardware Design Engineer with role in the...  ...of AWS next generation ML Chips, Cards and server... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $165.2k - $223.6k

    Annapurna Labs (our organization within AWS UC) designs silicon and software that accelerates innovation. Customers choose us to create...  ...supporting the ground-up development of key features that will support...  ...team members develop your engineering expertise so you feel... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  • $215k - $260k

     ...vertically integrated AI infrastructure...  ...owning the inference stack end to...  ...with customer engineering teams to...  ...prefill and decode disaggregation, request...  ...many kinds of ML models, with an...  ...and support the software and product features...  ...Professional development & tuition... 
    Temporary work

    Crusoe

    Sunnyvale, CA
    26 days ago
  • $200k - $322k

     ...Product Group builds AI solutions that...  ...Marketing Engineer focused on Enterprise AI Software, and accelerating...  ...AI blueprints, inference platforms,...  ...together across model development, inference, RAG...  ..., AI/ML, Data Science,...  ...GPU Operator, or Network Operator.Experience... 
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Software Development Engineer - AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs. Be the first to apply!