Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Large-Scale LLM Memory and Storage Systems

$272k - $431.25k

Jobleads-US

NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of any single GPU, this platform enables efficient, resilient deployment of cutting-edge LLM workloads. We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of large-scale LLM and storage systems. What you’ll be doing: Design and evolve a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage to support large-scale LLM inference. Architect and implement deep integrations with leading LLM serving engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters. Co-design interfaces and protocols that enable disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage (GPU, CPU, local disk, and remote memory) for high-throughput, low-latency inference. Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools. Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives). What we need to see: Masters or PhD or equivalent experience 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of delivering production services. Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency. Distributed caching or key-value systems, especially designs optimized for low latency and high concurrency. Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters. Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput. Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams. Ways to stand out from the crowd: Prior contributions to open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse. Experience designing unified memory or storage layers that expose a single logical KV or object model across GPU, host, SSD, and cloud tiers, especially in enterprise or hyperscale environments. Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML. With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our special engineering teams are growing fast. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA. #J-18808-Ljbffr

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Large-Scale LLM Memory and Storage Systems in Santa Clara, CA vacancy
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements...  ...GPU firmware and GPU system software, working...  ...firmware at fleet scale. You will drive work...  ...update orchestration for large-scale deployments —...  ...execution, compute kernels, memory hierarchy, and how... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements...  ...point for fleet-scale reliability, working...  ...enables you to distinguish systemic architectural gaps...  ...failure modes in large-scale GPU/accelerator...  ...compute, interconnect, memory, power, and thermal domainsExperience... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System...  ...above it, we treat memory safety, explicit ownership...  ...managing workloads at production scale. Drive triage of the...  ...discipline. Fluency in large, multi-repository codebases... 
    Suggested
    Full time
    Shift work

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...assistants and engineering-productivity...  ...we need a principal-level, hands-...  ...harden production systems and the...  ...ensure they scale.What you'll be...  ...like mature software, not prototypes...  ...agent workflows, memory and context...  ...patterns such as LLM-powered...  ...relevant to large-scale agent collaboration... 
    Suggested
    Full time
    Live in

    Nvidia

    Santa Clara, CA
    4 days ago
  • $174k - $252k

     ...test product or system development...  ...experience with software development in...  ....Experience in memory subsystem bring...  ...Google's software engineers develop the next...  ...at massive scale, and extend well...  ...distributed computing, large-scale system...  ...and data storage, security, artificial... 
    Suggested

    Google

    Mountain View, CA
    50 minutes ago
  • $272k - $431.25k

     ...organization seeks a Principal Engineer to architect and scale next-generation L10...  ...and L11 diagnostic systems for Cloud Service...  ...systems and hardware / software interfaces is...  ...systems, orchestrating large-scale stress...  ...GPUs, networking, memory, and high-speed interconnects... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $146.3k - $306.4k

    Defines architecture for large-scale systems software, firmware integration, and fleet automation that...  ..., and cost at hyperscale. Champions engineering excellence: coding standards, threat...  ...management paradigms across server, storage, networking, and GPU subsystems,... 
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    1 day ago
  • $114.6k - $234.6k

     ...Principal Systems Software Engineer Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve...  ...the intersection of software, firmware, hardware, and large-scale cloud infrastructure. You will work on complex GPU and... 
    Temporary work
    Flexible hours

    Hackajob

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...equipment manufacturers, and software providers to make...  ...facilities.We’re seeking a Principal Systems Software Engineer for Semiconductor Inspection...  ...real production latency, memory, throughput, reliability,...  ...crowd:Experience adapting large pretrained vision, multimodal... 
    Full time
    Local area
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...companies. As a Senior Principal Software Engineer at JPMorganChase within the...  ...root-cause acceleration, and large-scale refactoring/test...  ...architecting and deploying LLM & GNN solutions on AWS (e.g...  ...optimization and distributed systems for large models focused on... 

    JPMorgan Chase & Co.

    Palo Alto, CA
    a month ago
  • $272k - $431.25k

    NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation...  ...architecture, development, optimization, and scaling of foundational software...  ...optimization initiatives across CPU, GPU, memory, storage, networking, and platform subsystems... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...transformative areas of large language models and agentic systems, our mission is to...  ...of researchers and engineers who thrive on...  ...and queryability at scale. Develop production...  ...Architect context and memory systems for conversational...  ...Experience with LLM serving, agentic orchestration... 

    Boson AI

    Santa Clara, CA
    15 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU accelerated...  ...for accelerated computing to handle large data processing needs. Multi-node GPU...  ...problems challenges at large scale Provide recommendations and feedback... 
    Work experience placement

    Jobleads-US

    Santa Clara, CA
    4 days ago
  • $143k - $286k

     ...Home OfficePrincipal Software Engineer - Data Security PlatformAbout...  ...across Walmart’s systems, applications, and...  ...policy enforcement, large-scale data processing, and...  ...understanding of JVM internals, memory management,...  ...Deep understanding of storage engines, indexing,... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  •  ..., gaming and embedded systems. Grounded in a culture...  ...AMD is looking for a Principal-level PyTorch training...  ...scalability, and correctness of large-scale AI training on AMD...  ...performance, scaling, memory efficiency,...  ...communicate clearly with both engineers and stakeholders and... 

    AMD

    San Jose, CA
    8 hours ago
  • $220k - $250k

     ...StatesProducts - Engineering /Fulltime /HybridOver...  ...of position: Principal Software EngineerPosition...  ...: Director of SW Systems EngineeringApplication...  ...that process large volumes of network...  ...of large-scale distributed systems...  ...tool use, planning, memory, retrieval, orchestration... 
    Full time
    H1b
    Local area
    Work from home
    Work visa
    Shift work

    Extreme Networks, Inc.

    San Jose, CA
    4 days ago
  •  ...networking and storage are the critical...  ...unite these systems, making groundbreaking...  ...Infrastructure Engineering organization...  ...Staff Storage Software Engineer with...  ...protocol solutions at scale across object,...  ...of large-scale distributed...  ...technologies such as CXL memory pooling,... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...and highly motivated software professional to...  ...and Deep Learning Systems. As the complexity and scale of artificial intelligence...  ...utilization, memory bandwidth, cross-node...  ...Science, Computer Engineering, Electrical Engineering...  ...to, pioneering large language models.Research... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...cutting‑edge hardware and software innovation to deliver...  ...of forward‑thinking engineers tackling some of the...  ...searching for a Senior Systems Software Engineer with...  ...technical problems at large scale and help shape how AI...  ...architecture, networking, storage systems, and... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Corporation seeks a Production Storage Engineer to design, deploy, and optimize large-scale storage clusters for GPU-accelerated...  ...across distributed storage systems. You will work with cutting-edge...  ..., and collaborate with software, systems, and hardware teams to... 

    Jobleads-US

    Santa Clara, CA
    4 days ago
  • $193.13k - $257.5k

     ...up, operating at scale across Azure and AWS...  ...of agentic talent systems.What sets...  ...high standards. Our engineers, product leaders,...  ...will work with a large database of career...  ...a track record of software artifacts or academic...  ...optimization (vLLM, TensorRT-LLM).Desired Skills &... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    1 day ago
  •  ...Director, AI Systems Solutions Engineering Sunnyvale, CA About...  ...math, optimized scale-up networking, and memory architecture into...  ...purpose built for large-scale AI inference...  ...inference, serving software, and datacenter deployment...  ..., including LLM and multimodal... 
    Remote work

    Jobleads-US

    Sunnyvale, CA
    4 days ago
  • $183.6k - $297k

     ...TeamEngineering - Our engineering team is at the core of...  ...throughout the entire software development cycle. This...  ...tuning, and robust system architecture.Elevate Systems...  ...I/O, and designing large-scale, high-performance, or...  ...e.g., GitHub Copilot, LLM coding assistants) to... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $161k - $299k

     ...build high‑performance software engines and algorithms behind...  ...design productivity at scale. Role Overview...  ...distributed computing, and large‑scale numerical...  ...algorithmic efficiency, memory footprint, and multi‑core...  ...workloads across multi‑core systems and compute clusters... 

    Jobleads-US

    San Jose, CA
    4 days ago
  •  ...looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure...  ...of machine learning systems, including large language models,...  ...models run reliably at scale. This role is a strong...  ...model throughput, latency, memory efficiency, scalability... 

    Jobleads-US

    San Jose, CA
    4 days ago
  • $226k - $369k

     ...our world-class software engineering team, you will take...  ..., scalable data storage infrastructure, graph...  ..., API design and systems design, and your...  ...at massive scale. LinkedIn has pioneered...  ...our company.As a Principal Staff Software...  ...developing and scaling large scale databases... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  • $224k - $356.5k

     ...highly motivated, creative engineers to join the Platform Software team. You will work...  ...aspects of SOC and system, and technology verticals...  ...that apply to large complex systems deployed at scale.Strong C/C++ and Python...  ...fundamentals (caches, buses, memory controllers, DMA, etc.... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $190k - $237k

     ...unmanned aircraft systems (“UAS”), aviation-...  ...Staff AI Systems Engineer, you will architect...  ...services required for large-scale AI model training...  ...advanced LLM serving engines.Cross...  ...AI researchers and Software Engineers to productionize...  ..., and columnar storage formats optimized... 
    Local area
    Worldwide
    Visa sponsorship

    Archer Aviation

    San Jose, CA
    8 hours ago
  •  ...interconnect, memory and boot path,...  ...IP, firmware, software and field teams...  ...a passion for system architecture, silicon...  ...with engineers located in different...  ...5/6), boot and storage (OSPI, UFS, SD/...  ...Workflows: Use AI/LLM/ML tools to surface...  ...and correlate large test and log... 

    AMD

    San Jose, CA
    3 days ago
  • $165.2k - $223.6k

     ...AWS Neuron, the software development kit...  ...boundary, our engineers build systematic...  ...that are very large, yet our teams...  ...performance at scale for customers and...  ...variety of LLM model families,...  ...the stack from system level optimizations...  ..., optimizing memory usage, and shaping... 
    Full time
    Work experience placement
    Internship
    Local area
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 hour ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Large-Scale LLM Memory and Storage Systems. Be the first to apply!