Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Large-Scale LLM Memory and Storage Systems

$272k - $431.25k

NVIDIA

NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of any single GPU, this platform enables efficient, resilient deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of large-scale LLM and storage systems.What you'll be doing:Design and evolve a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage to support large-scale LLM inference.Architect and implement deep integrations with leading LLM serving engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters.Co-design interfaces and protocols that enable disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage (GPU, CPU, local disk, and remote memory) for high-throughput, low-latency inference.Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools.Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives).What we need to see:Masters or PhD or equivalent experience15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of delivering production services.Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency.Distributed caching or key-value systems, especially designs optimized for low latency and high concurrency.Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters.Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput.Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams.Ways to stand out from the crowd:Prior contributions to open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse.Experience designing unified memory or storage layers that expose a single logical KV or object model across GPU, host, SSD, and cloud tiers, especially in enterprise or hyperscale environments.Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML.With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our special engineering teams are growing fast. If you're a creative and autonomous engineer with a genuine passion for technology, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until January 13, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, Remote; US, MA, RemoteType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Large-Scale LLM Memory and Storage Systems in Washington DC vacancy
  •  ...Principal Software Engineer Mastercard is a global technology...  ...combine enterprise-scale technical leadership...  ...Design and implement large scale distributed systems. Develop...  ...distributed caches, in-memory data grids. Real...  ...Agentic AI patterns, LLM integration, and prompt... 
    Suggested
    Remote work
    Work from home

    Dynamic Yield

    Arlington, VA
    2 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference)...  ...and reliable AI systems, changing banking...  ..., and support AI software components including...  ...model training, large language model inference...  ...throughput - of large scale production AI...  ...VectorDBs, Guardrails, Memory) using Python,... 
    Suggested
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    McLean, VA
    16 hours ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (LLM Customization and Finetuning...  ...and reliable AI systems, changing banking...  ..., and support AI software components...  ...foundation model training, large language model...  ...throughput - of large scale production AI...  ...VectorDBs, Guardrails, Memory) using Python, C++... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    1 day ago
  •  ...missions and modernize their systems, so that they can...  ...skilled AI Systems Engineer who will be part of a...  ...MLOps practices. The Software Engineer will work closely...  ...* Proficiency with LLM development practices...  ...(Azure, AWS, GCP) or large scale data ecosystems. * Understanding... 
    Suggested

    Tria Federal

    Arlington, VA
    2 days ago
  • $197.3k - $225.1k

    Lead AI Engineer (AI Foundations, LLM Core and Agentic AI) Overview...  ...and reliable AI systems, changing banking...  ..., and support AI software components...  ...model training, large language model inference...  ...— of large scale production AI systems...  ...VectorDBs, Guardrails, Memory) using Python,... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    McLean, VA
    4 days ago
  • $172.5k - $313.7k

     ...Salesforce.Salesforce is seeking a Principal Software Engineer (PMTS) to join the Tableau...  ...business processes at scale.This role sits at the intersection of large-scale distributed systems, data-intensive analytics,...  ...database, cache, and storage layers.Expertise in shaping... 
    Full time

    Salesforce

    Washington DC
    2 days ago
  • $163.8k - $257.4k

     ...happen-fast.About the Role:We are seeking an experienced Principal Software Engineer to lead the design and development of ZoomInfo's...  ...grained authorization, rate limiting.- Familiarity with large-scale data systems; exposure to BigTable, BigQuery, or Solr is a plus.- Experience... 
    Worldwide

    ZoomInfo

    Bethesda, MD
    1 day ago
  • $13 per hour

     ...is now hardening and scaling its foundations as customer...  ..., and scaling the systems that connect our...  ...experienced Senior, Lead, and Principal Engineers to serve as a key...  ...production-grade software using modern engineering...  ...and operating large-scale, user-facing web... 
    Full time

    Salesforce

    Washington DC
    12 hours ago
  • $145k - $200k

     ...builds the world’s leading software for data-driven...  ...seeking a Senior Software Engineer to join a customer-facing...  ...for Autonomous Systems for use in operational...  ...threat detectionBuild LLM and AI-driven systems capable...  ...patterns, reliability and scaling) of new and existing systems... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  •  ...done. We’re hiring a seasoned Senior Software Engineer to help us build a world class developer...  ...team within ES (Engineering Systems), you will play a central role in setting...  ...infrastructure and tooling to be elastic, large-scale, and highly performant with simplicity... 
    Full time

    Snowflake

    Washington DC
    16 hours ago
  •  ...Design and develop software applications using Twelve...  ...mainframe and Linux operating systems and other platforms,...  ...Science, Mathematics, Engineering or a related field....  ...capable of handling large volumes of concurrent...  ...applications for effective CPU, Memory utilization, Quick... 

    Saxon Global

    Riverdale, MD
    1 day ago
  • $144.2k - $288.4k

     ...Principal Software Engineer We're building a world of health around every individual...  ...support applications. These systems combine low-latency...  ...architectures with advanced LLM, OCR, and ML pipelines running...  ...delivering production systems at scale ~5+ years as a hands-on... 
    Hourly pay
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Oak St. Health

    Washington DC
    3 days ago
  • $114.6k - $234.6k

     ...offers unmatched hyper-scale, multi-tenant...  ...Frameworks: Spearhead the engineering of new container...  ...persistent storage and networking solutions...  ...initiatives with large-scale customer and...  ...at-scale system and data-plane architectures...  ...and drive the software design and... 
    Temporary work
    Work experience placement
    Worldwide
    Flexible hours

    Oracle

    Washington DC
    5 days ago
  • $150k - $175k

     ...looking for an experienced Principal Software Engineer who will be responsible for...  ...using source code control system and deployment pipelines...  ...production software as necessary Scale and tune for performance to...  ...AI Training & Inference: LLM (Large Language Models), Deep... 
    Work at office

    Kastle Systems

    Falls Church, VA
    1 day ago
  • $197.3k - $313.7k

     ...investing in a unified engineering ecosystem that spans...  ...motivated, hands-on Principal Member of Technical Staff...  ...rare combination of software engineering, data...  ...while building scalable systems that connect...  ....Architect and build large-scale applications using Salesforce... 
    Full time

    Salesforce

    Washington DC
    3 days ago
  • $163.8k - $257.4k

     ...seeking an experienced Principal Software Engineer (Data) to join our...  ...class data pipelines and systems to process billions...  ...architectures and storage solutions to efficiently...  ...building and scaling big data pipelines and...  ...Experience working with very large datasets and... 
    Worldwide

    ZoomInfo

    Bethesda, MD
    2 days ago
  • $144.2k - $288.4k

     ...on, passionate, and people-focused Principal Software Development Engineer to join our high-energy, growing team...  ..., you will operate within a large-scale product, engineering, and architecture...  ...an Agile environment. Exceptional systems thinker who can seamlessly navigate... 
    Hourly pay
    Full time
    Temporary work
    Local area
    Flexible hours

    Stryker

    Washington DC
    4 days ago
  •  ...stronger economic outcomes at scale through harnessing the...  ...with health systems across the U.S., we have...  ...hiring a product-minded AI Engineer to help build and improve...  ...~2–3 years of software engineering experience...  ...Exposure to AI/ML or LLM-based applications Experience... 
    Full time

    Sage Care

    Washington DC
    16 hours ago
  • $117.2k - $313.7k

     ...and you are the future of Salesforce.Distributed Systems Software Engineer - Public Cloud (Senior/Lead/Principal) Note: By applying to the Public Cloud -...  ...are responsible for innovating and maintaining a large scale distributed systems engineering platform that ships... 
    Full time

    Salesforce

    Washington DC
    3 days ago
  • $152k - $241.5k

    NVIDIA is hiring experienced software engineers to help scale up its AI Infrastructure. We expect you to have significant software engineering...  ...of an DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI... 
    Full time
    Remote work

    Nvidia

    Washington DC
    4 days ago
  • $196k - $294k

     ...ABOUT THE TEAM ~ We are seeking a Principal Software Engineer to lead architectural and strategic...  ...who quickly understands distributed systems and creates features that drive bottom...  ...platforms. ~ Background in leading large-scale CI/CD and testing frameworks. ~ Familiarity... 
    Full time
    Work experience placement
    Local area
    Relocation package

    Anduril Industries

    Washington DC
    16 hours ago
  • $125k - $140k

     ...highly skilled and self-driven Lead Virtualization Systems Engineer (Converged) to support large-scale enterprise datacenters encompassing over 3,000 virtual...  ...a VMware-based infrastructure and associated storage systems in a government environment. This senior-level... 
    Full time
    Contract work
    Flexible hours

    Seneca Holdings LLC

    Beltsville, MD
    2 days ago
  • $148.5k - $313.7k

     ...By applying to the Senior / Lead / Principal Software Engineer - Foundations Team posting,...  ...the pace of a startup backed by the scale and stability of a global enterprise...  ...driving architecture decisions, solving large-scale distributed systems challenges, and building... 
    Full time
    Work experience placement
    Remote work

    Salesforce

    Washington DC
    2 days ago
  • $157k - $224k

     ...STR is hiring a Lead Model-Based Systems Engineer (MBSE) in our Woburn, MA office to work across...  ...and the ability to work with complex software systems and the utilization of Digital...  ...architectures and integrating with large scale software-intensive systems and platform... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Flexible hours

    Systemstechnologyresearch

    Arlington, VA
    1 day ago
  • $75 - $80 per hour

     ...Recruiter at Matlen Silver Job Title: Senior AI/LLM Engineer Duration: 12+ Months Location: Washington, DC Required Pay Scale: $75-$80/hour W2 ***Due to client requirements...  ...advanced AI solutions with a strong focus on Large Language Models (LLMs). The ideal candidate... 
    Full time

    Matlen Silver

    Washington DC
    4 days ago
  •  ...Senior Software Engineer – Platform Systems II Confidential Client – Purple Squirrel Enterprises About the Opportunity Our client is a...  ..., Rust, or Java ~ Experience designing and operating large-scale, distributed, or network-aware systems ~ Professional... 
    Full time
    Flexible hours

    Purple Squirrel Enterprises

    Washington DC
    16 hours ago
  • $168.1k - $227.4k

    As a Software Development Engineer for Amazon eXperiences & Upskilling (...  ...technical expertise in large language models with...  ...intelligent agent systems.In this highly strategic...  ...that leverage LLM & SLM technologies. You...  ...patterns, reliability and scaling) of new and existing... 
    Internship
    Flexible hours

    Amazon

    Arlington, VA
    12 hours ago
  • $143.7k - $194.4k

     ...and more! Our Echo Software team works on not only...  ...Devices.- Driving engineering best practices (e.g....  ...patterns, reliability and scaling) of new and existing systems experience- 1+ years...  ...and developing large-scale, multi-tiered,...  ...Machine Learning and LLM fundamentals,... 
    Internship
    Immediate start
    Worldwide
    Flexible hours

    Amazon

    Arlington, VA
    4 days ago
  • $200k - $280k

     ...clearance.*** Are you a Principal Full Stack Software Engineer who is ready for a...  ...on Software & System Engineering in Enterprise...  ...and systems to a large government contract....  ...infrastructure and scale future capabilities....  ...configurations like Skills and Memory.   Familiarity... 
    Full time
    Contract work
    Remote work
    Work from home
    Relocation package

    GliaCell Technologies

    Laurel, MD
    16 days ago
  • $320k - $405k

     ...and steerable AI systems. We want AI to be...  ...committed researchers, engineers, policy experts,...  ...performance, memory usage, and startup...  ...of experience as a software engineer, with strong...  ...re working with a large number of technologies...  ...just a few large-scale research efforts.... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Washington DC
    16 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Large-Scale LLM Memory and Storage Systems. Be the first to apply!