Principal Software Engineer - Large-Scale LLM Memory and Storage Systems
$272k - $425.5kNVIDIA
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems page is loaded## Principal Software Engineer – Large-Scale LLM Memory and Storage Systemslocations: US, CA, Santa Clara: US, WA, Remote: US, MA, Remotetime type: Full timeposted on: Posted Todayjob requisition id: JR2010271NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of any single GPU, this platform enables efficient, resilient deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of large-scale LLM and storage systems.**What you'll be doing:*** Design and evolve a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage to support large-scale LLM inference.* Architect and implement deep integrations with leading LLM serving engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters.* Co-design interfaces and protocols that enable disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage (GPU, CPU, local disk, and remote memory) for high-throughput, low-latency inference.* Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools.* Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives).**What we need to see:*** Masters or PhD or equivalent experience* 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of delivering production services.* Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency.* Distributed caching or key-value systems, especially designs optimized for low latency and high concurrency.* Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters.* Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput.* Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams.**Ways to stand out from the crowd:*** Prior contributions to open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse.* Experience designing unified memory or storage layers that expose a single logical KV or object model across GPU, host, SSD, and cloud tiers, especially in enterprise or hyperscale environments.* Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML.With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our special engineering teams are growing fast. If you're a creative and autonomous engineer with a genuine passion for technology, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 425,500 USD.You will also be eligible for equity and .Applications for this job will be accepted at least until December 26, 2025.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr
- ...transformative areas of large language models and agentic systems, our mission is to... ...of researchers and engineers who thrive on... ...and queryability at scale Develop production‑grade... ...Architect context and memory systems for conversational... ...Experience with LLM serving, agentic...Suggested
$197.3k - $225.1k
...Lead AI Engineer (AI Foundations, LLM Customization and Finetuning... ...and reliable AI systems, changing banking... ..., and support AI software components... ...model training, large language model inference... ...throughput — of large scale production AI... ..., Guardrails, Memory) using Python, C++...SuggestedFull timePart timeLocal area$197.3k - $225.1k
Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital... ...and reliable AI systems, changing banking... ..., and support AI software components... ...model training, large language model inference... ...— of large scale production AI systems... ..., Guardrails, Memory) using Python, C++...SuggestedFull timePart timeLocal area$197.3k - $225.1k
...Lead AI Engineer (AI Foundations, LLM Core and Agentic AI) Overview... ...and reliable AI systems, changing banking... ..., and support AI software components... ...model training, large language model inference... ...— of large scale production AI systems... ..., Guardrails, Memory) using Python, C++...SuggestedFull timePart timeLocal area$272k - $431.25k
...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation... ...architecture, development, optimization, and scaling of foundational software... ...optimization initiatives across CPU, GPU, memory, storage, networking, and platform subsystems...Suggested- ...AI agents and agentic systems. Architect and implement... ...Develop and integrate large language models (LLMs)... ...operation of AI agents at scale. Diagnose and... ...proven by a track record of software artifacts or academic publications... ...(vLLM, TensorRT-LLM). Research experience in...Full timeWork experience placement
- ...intelligence. Its robots are engineered to perform a variety... ...to create embodied AI systems that can perceive the... ...pixels, reason over memory, and reliably execute... ...ML systems Strong software engineering skills and... ...Experience working with large-scale distributed training...Full time
$349k - $431k
...states. Waymo's Systems Intelligence and... ...autonomous driving software. Waymo's AI is at... ...increasingly leveraging large-scale Foundation Models... ...our Director of Engineering who leads Systems... ...experienced Principal Software Engineer... ...data pipelines, storage systems, and tokenization...- ...Description We are hiring a Principal Engineer to serve as an... ...for agentic systems, MCP ecosystem, LLM gateway, memory and knowledge layers... ...the platform scales securely and reliably... ...years of professional software engineering... ...designing and operating large‑scale AI/ML,...Temporary workRemote workFlexible hoursShift work
$175k - $263k
...fundamentally reshaping the data storage industry. Here, you... ...found in unit testing, system testing and customer... ...to operate and scale in a distributed system... ...and production quality software ~ Proven design sensibility... ...Named Fortune's Best Large Workplaces in the Bay Area...Full timeWork at officeFlexible hours$175k - $317k
...Join to apply for the Software Engineer role at Pure Storage Get AI-powered advice on this... ...algorithms and high-performance system designs. You will drive... ...and debugging at scale. Work in an in-office environment... ...been named Fortune's Best Large Workplaces in the Bay Area...Work at officeFlexible hours- ...testing framework and library. He/she will participate in the core system design and development. Our target system is based on jepsen... ...~5+ years proven records on infrastructure level software development experience ~2+ years Clojure development experience...Full time
$132.3k - $198.45k
...path to AVs at commercial scale, empowering a safer, richer... ...We’re looking for senior engineers to build/scale Nuro's large-scale computing infrastructure... ...cloud/data center. This system is the foundation of many... ..., e.g. CPU scheduler, memory management, file systems...Full time$140k - $210k
...Yubico is looking for a Senior Software Engineer who is innovative and has a... ...production software to enhance and scale our capabilities and ease of... ...Familiarity with operating system internals and services, such... ...software using native and/or memory safe languages such as modern...Full timeContract workWork at officeImmediate startRemote work$231.4k - $331.8k
...s Platform & Identity Engineering Group, a foundational... ...responsible for building and scaling the core platform... ...the underlying systems that support Cisco's innovation... ...direction of Cisco's software and technology... ...designing and developing large-scale software, platform...Full timeTemporary workLocal areaFlexible hours$221.2k - $387.1k
...It all started when engineer Fred Luddy wrote code that automated... ...people. Job Description Principal Software Engineer The engineering... ...operations at enterprise scale. Together, these... ...era. The work will span large-scale distributed systems, data ingestion, enterprise...Work at officeImmediate startRemote workFlexible hours$99.6k - $234.6k
...Principal Software Development Engineer As a Principal Software Development Engineer in... ...integration to distributed systems architecture and operational excellence in large-scale production environments.... ...services, such as Database, Storage Services, SaaS, and other...Temporary workFlexible hours$208k - $260k
...governments and educational organizations. As a Principal Software Engineer on the Network Management System team, you will lead the design and development of... ...extensible, high-performance platforms that support large-scale deployments and long-term product evolution...Local areaWorldwide3 days per week$196.8k - $262.7k
...highly skilled Senior AI Engineer to design and implement advanced agentic AI systems. In this role, you'll... ...pipelines that combine LLM‑based reasoning, knowledge... ...of RAG pipelines, memory management, and evaluation... ...and operations in large‑scale production environments...Immediate startRemote workVisa sponsorship$125k - $191.7k
...categorized as hybrid/Remote. Role: As a Senior Software Systems Engineer on the Software Validation team within... ...that support continuous and scaled software release cycles. Combine experience... ...SQL, Python, and C++ for analyzing large data sets. Strong analytical thinking...Local areaRemote workFlexible hours$249k - $348.5k
Principal Software Development Engineer Our Technology Team partners with teams across... ...failure domains that scale horizontally with... ...building, and operating large‑scale, cloud‑native distributed systems and platform services... ...compute, networking, storage, service discovery,...Flexible hours- ...visionary Senior Principal Software Engineer / Technical Staff... ...technical debt, and scaling global... ...build, and deploy AI/LLM-driven applications (e.g., RAG systems, log analyzers, agentic... ...-leading large-scale operations... ...customizable computing, storage and networking solutions...Temporary workFor contractorsWork at officeLocal areaWorldwideShift work
$227k - $303k
...innovators to build and scale AI with... ...the services, systems, and developer-... ...that enable every engineer at CoreWeave to... ...build and ship software faster, more... ...Role This is a principal-level software... ...who has built large-scale production... ...that integrate LLM-driven workflows...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hours$176.1k - $308.2k
...Staff AI Engineer - Conversational & Agentic... ...coordination, memory management, and... ...by the global scale of ServiceNow... ...management systems — versioning, templating... ...decisions LLM integration :... ...building production software systems with a... ...depth in how large language models...Full timeWork at officeImmediate startRemote workFlexible hours- ...and reliable AI systems, changing... ...applied science and engineering teams to deliver... ...and support AI software components... ...model training, large language model... ...state‑of‑the‑art LLM optimization techniques... ...—of large‑scale production AI... ...VectorDBs, guardrails, memory) using Python,...Local area
$197.3k - $225.1k
...Lead AI Engineer (MLX, Agentic AI, Gen... ...and reliable AI systems, changing banking... ...and support AI software components... ...model training, large language model... ...state‑of‑the‑art LLM optimization techniques... ...—of large scale production AI... ...Vectordb, guardrails, memory) using Python,...Full timePart timeLocal area$229.9k - $262.4k
...Capital One, you’ll join a large group of makers, breakers, doers... ...seeking Full Stack/Back End Software Engineers who are passionate about... ...deployment of AI/ML solutions at scale. MLX Tech harnesses the... ...microservices, and full stack systems to create solutions that...InternshipLocal area$313.06k
...brand by hundreds of large institutions... ...We are seeking a Principal Engineer with a deep expertise... ...intelligent agent systems on our global crypto... ...in AGI, LLM-integrated agents,... ...context engineering, memory management, tool use... ...fast-paced, high-scale environments (preferably...$100k - $150k
...AI Systems Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data... ...platform layer that powers large-scale AI training and inference... ...frameworks, scheduling, storage performance, and developer...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$192.1k - $249.6k
...or the EL7), a mid-large five-seater smart... ...hybrid inference systems, dynamically distributing... ...Language Model (LLM) and Vision-... ...team of specialized engineers working across distributed... ...: Designing, scaling, and deploying... ...FlashAttention), and advanced memory management (e.g.,...Full timeTemporary workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer - Large-Scale LLM Memory and Storage Systems. Be the first to apply!
- principal software engineer Santa Clara, CA
- senior principal scientist Santa Clara, CA
- principal Santa Clara, CA
- senior principal cloud computing engineer Santa Clara, CA
- software engineer - cloud services Santa Clara, CA
- software sales executive Santa Clara, CA
- entry level software sales Santa Clara, CA
- healthcare software sales Santa Clara, CA
- software implementation project manager Santa Clara, CA
- internship software Santa Clara, CA


