Principal Software Engineer - Large-Scale LLM Memory and Storage Systems
$272k - $431.25kJobleads-US
NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of any single GPU, this platform enables efficient, resilient deployment of cutting-edge LLM workloads. We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of large-scale LLM and storage systems. What you’ll be doing: Design and evolve a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage to support large-scale LLM inference. Architect and implement deep integrations with leading LLM serving engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters. Co-design interfaces and protocols that enable disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage (GPU, CPU, local disk, and remote memory) for high-throughput, low-latency inference. Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools. Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives). What we need to see: Masters or PhD or equivalent experience 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of delivering production services. Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency. Distributed caching or key-value systems, especially designs optimized for low latency and high concurrency. Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters. Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput. Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams. Ways to stand out from the crowd: Prior contributions to open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse. Experience designing unified memory or storage layers that expose a single logical KV or object model across GPU, host, SSD, and cloud tiers, especially in enterprise or hyperscale environments. Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML. With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our special engineering teams are growing fast. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA. #J-18808-Ljbffr
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements... ...GPU firmware and GPU system software, working... ...firmware at fleet scale. You will drive work... ...update orchestration for large-scale deployments —... ...execution, compute kernels, memory hierarchy, and how...SuggestedFull timeRemote work$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements... ...point for fleet-scale reliability, working... ...enables you to distinguish systemic architectural gaps... ...failure modes in large-scale GPU/accelerator... ...compute, interconnect, memory, power, and thermal domainsExperience...SuggestedFull time$272k - $431.25k
...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System... ...above it, we treat memory safety, explicit ownership... ...managing workloads at production scale. Drive triage of the... ...discipline. Fluency in large, multi-repository codebases...SuggestedFull timeShift work$272k - $431.25k
...assistants and engineering-productivity... ...we need a principal-level, hands-... ...harden production systems and the... ...ensure they scale.What you'll be... ...like mature software, not prototypes... ...agent workflows, memory and context... ...patterns such as LLM-powered... ...relevant to large-scale agent collaboration...SuggestedFull timeLive in$174k - $252k
...test product or system development... ...experience with software development in... ....Experience in memory subsystem bring... ...Google's software engineers develop the next... ...at massive scale, and extend well... ...distributed computing, large-scale system... ...and data storage, security, artificial...Suggested$272k - $431.25k
...organization seeks a Principal Engineer to architect and scale next-generation L10... ...and L11 diagnostic systems for Cloud Service... ...systems and hardware / software interfaces is... ...systems, orchestrating large-scale stress... ...GPUs, networking, memory, and high-speed interconnects...Full time$146.3k - $306.4k
Defines architecture for large-scale systems software, firmware integration, and fleet automation that... ..., and cost at hyperscale. Champions engineering excellence: coding standards, threat... ...management paradigms across server, storage, networking, and GPU subsystems,...Temporary workFlexible hoursShift work$114.6k - $234.6k
...Principal Systems Software Engineer Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve... ...the intersection of software, firmware, hardware, and large-scale cloud infrastructure. You will work on complex GPU and...Temporary workFlexible hours$272k - $431.25k
...equipment manufacturers, and software providers to make... ...facilities.We’re seeking a Principal Systems Software Engineer for Semiconductor Inspection... ...real production latency, memory, throughput, reliability,... ...crowd:Experience adapting large pretrained vision, multimodal...Full timeLocal areaShift work- ...companies. As a Senior Principal Software Engineer at JPMorganChase within the... ...root-cause acceleration, and large-scale refactoring/test... ...architecting and deploying LLM & GNN solutions on AWS (e.g... ...optimization and distributed systems for large models focused on...
$272k - $431.25k
NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation... ...architecture, development, optimization, and scaling of foundational software... ...optimization initiatives across CPU, GPU, memory, storage, networking, and platform subsystems...Full time- ...transformative areas of large language models and agentic systems, our mission is to... ...of researchers and engineers who thrive on... ...and queryability at scale. Develop production... ...Architect context and memory systems for conversational... ...Experience with LLM serving, agentic orchestration...
$272k - $431.25k
...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU accelerated... ...for accelerated computing to handle large data processing needs. Multi-node GPU... ...problems challenges at large scale Provide recommendations and feedback...Work experience placement$143k - $286k
...Home OfficePrincipal Software Engineer - Data Security PlatformAbout... ...across Walmart’s systems, applications, and... ...policy enforcement, large-scale data processing, and... ...understanding of JVM internals, memory management,... ...Deep understanding of storage engines, indexing,...Full timeTemporary workPart time- ..., gaming and embedded systems. Grounded in a culture... ...AMD is looking for a Principal-level PyTorch training... ...scalability, and correctness of large-scale AI training on AMD... ...performance, scaling, memory efficiency,... ...communicate clearly with both engineers and stakeholders and...
$220k - $250k
...StatesProducts - Engineering /Fulltime /HybridOver... ...of position: Principal Software EngineerPosition... ...: Director of SW Systems EngineeringApplication... ...that process large volumes of network... ...of large-scale distributed systems... ...tool use, planning, memory, retrieval, orchestration...Full timeH1bLocal areaWork from homeWork visaShift work- ...networking and storage are the critical... ...unite these systems, making groundbreaking... ...Infrastructure Engineering organization... ...Staff Storage Software Engineer with... ...protocol solutions at scale across object,... ...of large-scale distributed... ...technologies such as CXL memory pooling,...Work at officeLocal areaWork from homeFlexible hours
$184k - $287.5k
...and highly motivated software professional to... ...and Deep Learning Systems. As the complexity and scale of artificial intelligence... ...utilization, memory bandwidth, cross-node... ...Science, Computer Engineering, Electrical Engineering... ...to, pioneering large language models.Research...Full time$184k - $287.5k
...cutting‑edge hardware and software innovation to deliver... ...of forward‑thinking engineers tackling some of the... ...searching for a Senior Systems Software Engineer with... ...technical problems at large scale and help shape how AI... ...architecture, networking, storage systems, and...Full timeRemote work- ...Corporation seeks a Production Storage Engineer to design, deploy, and optimize large-scale storage clusters for GPU-accelerated... ...across distributed storage systems. You will work with cutting-edge... ..., and collaborate with software, systems, and hardware teams to...
$193.13k - $257.5k
...up, operating at scale across Azure and AWS... ...of agentic talent systems.What sets... ...high standards. Our engineers, product leaders,... ...will work with a large database of career... ...a track record of software artifacts or academic... ...optimization (vLLM, TensorRT-LLM).Desired Skills &...Work experience placementWork at officeRemote workFlexible hours3 days per week- ...Director, AI Systems Solutions Engineering Sunnyvale, CA About... ...math, optimized scale-up networking, and memory architecture into... ...purpose built for large-scale AI inference... ...inference, serving software, and datacenter deployment... ..., including LLM and multimodal...Remote work
$183.6k - $297k
...TeamEngineering - Our engineering team is at the core of... ...throughout the entire software development cycle. This... ...tuning, and robust system architecture.Elevate Systems... ...I/O, and designing large-scale, high-performance, or... ...e.g., GitHub Copilot, LLM coding assistants) to...Full timeWork at office$161k - $299k
...build high‑performance software engines and algorithms behind... ...design productivity at scale. Role Overview... ...distributed computing, and large‑scale numerical... ...algorithmic efficiency, memory footprint, and multi‑core... ...workloads across multi‑core systems and compute clusters...- ...looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure... ...of machine learning systems, including large language models,... ...models run reliably at scale. This role is a strong... ...model throughput, latency, memory efficiency, scalability...
$226k - $369k
...our world-class software engineering team, you will take... ..., scalable data storage infrastructure, graph... ..., API design and systems design, and your... ...at massive scale. LinkedIn has pioneered... ...our company.As a Principal Staff Software... ...developing and scaling large scale databases...For contractorsWork at officeFlexible hours$224k - $356.5k
...highly motivated, creative engineers to join the Platform Software team. You will work... ...aspects of SOC and system, and technology verticals... ...that apply to large complex systems deployed at scale.Strong C/C++ and Python... ...fundamentals (caches, buses, memory controllers, DMA, etc....Full time$190k - $237k
...unmanned aircraft systems (“UAS”), aviation-... ...Staff AI Systems Engineer, you will architect... ...services required for large-scale AI model training... ...advanced LLM serving engines.Cross... ...AI researchers and Software Engineers to productionize... ..., and columnar storage formats optimized...Local areaWorldwideVisa sponsorship- ...interconnect, memory and boot path,... ...IP, firmware, software and field teams... ...a passion for system architecture, silicon... ...with engineers located in different... ...5/6), boot and storage (OSPI, UFS, SD/... ...Workflows: Use AI/LLM/ML tools to surface... ...and correlate large test and log...
$165.2k - $223.6k
...AWS Neuron, the software development kit... ...boundary, our engineers build systematic... ...that are very large, yet our teams... ...performance at scale for customers and... ...variety of LLM model families,... ...the stack from system level optimizations... ..., optimizing memory usage, and shaping...Full timeWork experience placementInternshipLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer - Large-Scale LLM Memory and Storage Systems. Be the first to apply!
- senior principal software engineer Santa Clara, CA
- principal software engineer Santa Clara, CA
- principal cloud computing engineer Santa Clara, CA
- senior principal cloud computing engineer Santa Clara, CA
- principal architect Santa Clara, CA
- principal Santa Clara, CA
- senior principal scientist Santa Clara, CA
- ultimate software Santa Clara, CA
- software qa Santa Clara, CA
- embedded software Santa Clara, CA



