Principal Software Engineer - Large-Scale LLM Memory and Storage Systems
$272k - $431.25kNVIDIA
NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across heterogeneous clusters so that many accelerators feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of any single GPU, this platform enables efficient, resilient deployment of cutting-edge LLM workloads. We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of large-scale LLM and storage systems. What You'll Be Doing Design and evolve a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage to support large-scale LLM inference. Architect and implement deep integrations with leading LLM serving engines (such as vLLM, SGLang, TensorRT-LLM), with a focus on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters. Co-design interfaces and protocols that enable disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage (GPU, CPU, local disk, and remote memory) for high-throughput, low-latency inference. Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools. Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives). What We Need To See Masters or PhD or equivalent experience. 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of delivering production services. Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency. Distributed caching or key-value systems, especially designs optimized for low latency and high concurrency. Hands‑on experience with networked I/O and RDMA/NVMe‑oF/NVLink‑style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters. Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput. Excellent communication skills and prior experience leading cross‑functional efforts with research, product, and customer teams. Ways To Stand Out From The Crowd Prior contributions to open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse. Experience designing unified memory or storage layers that expose a single logical KV or object model across GPU, host, SSD, and cloud tiers, especially in enterprise or hyperscale environments. Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML. With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until January 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
JR2010271
#J-18808-Ljbffr NVIDIA$272k - $431.25k
...feel like a single system at datacenter scale. As large language models rapidly outgrow the memory and compute budget of... ...deployment of cutting-edge LLM workloads. We are seeking a Principal Systems Engineer to define the vision... ...large-scale LLM and storage systems. What You'll...SuggestedLocal areaRemote work$154.6k - $180k
...globally uses our network, scale, connectivity and... ...The hybrid‑remote Principal Software Development Engineer leads the design,... ...of complex software systems and applications. This... ...Model Optimization (LLM & SLM): Partner with... .../extraction and Large Language Models (LLMs...SuggestedLocal areaRemote work- Advanced Monitored Caregiving Inc. is seeking a Senior Software Engineer for Applied AI focused on voice agents and ML systems. You will work across real-time streaming,... ...guardrails and cost‑aware architectures for LLM workloads, ship low-latency code, and own end‑to...Suggested
$144.2k - $288.4k
...Position Summary We are hiring a Principal Software Engineer — a deeply hands‑on... ...support applications. These systems combine low‑latency distributed... ...architectures with advanced LLM, OCR, and ML pipelines running... ...production systems at scale 5+ years as a hands‑on Staff...SuggestedHourly payFull timeTemporary workLocal areaFlexible hours$116.3k - $174.5k
...opportunities to work on revolutionary systems that impact people's lives around the... ...Grumman is seeking an experienced Sr Principal Software Engineer to join our dynamic team. This... ...Responsibilities Leadership and standup of large-scale software factory used by software...SuggestedRelocation packageShift work- ...every day. CMT is seeking a Principal Backend Software Engineer to produce scalable... ...and evolution of the backend systems that power DriveWell Fusion. You will design and scale high-throughput data pipelines... ..., feature engineering, and large-scale inference — including...
- ...job Tech Lead, Data & Inference Engineer Our Client About Us Catalyst... ...evolving world of intelligent systems. Work type : Full Time, Lead the design, development and scaling of an end to end data platform... ...scoring and quality assurance using large language models and retrieval...Full time
$135.2k - $306.4k
...offers unmatched hyper-scale, multi-tenant... ...Frameworks: Spearhead the engineering of new container... ...persistent storage and networking solutions... ...initiatives with large-scale customer and... ...at-scale system and data-plane architectures... ...and drive the software design and...Temporary workWork experience placementWorldwideFlexible hours- ...businesses connect and scale on premises, in the... ...across Teradata’s memory and execution systems, ensuring correctness... ...under complex, large‑scale workloads. You... ...wide issues Influence engineering decisions through deep... ...PDE engineering teams Storage, NOS, and execution‑...Permanent employmentFlexible hours
$79k - $119k
...Relativity’s legal AI software to securely... ...Do At Relativity, engineers don’t just write code... ...code. They build the systems that power AI-... ...workflows at massive scale, enabling... ...systems that handle large-scale AI traffic and... ...large language model (LLM) APIs or...Remote workHome office- EY is seeking a Delivery Lead to plan, drive, and manage large SAP implementations. You will establish operating structures for large programs and ensure day-to-day delivery leads to successful client outcomes. The role requires leading engagements, budgeting, and ensuring...
$75.1k - $118.9k
...of the Time Overview Northrop Grumman Aeronautics Systems has openings for Software Engineer / Principal Software Engineer - Simulation to join our team of... ...understanding of C and C++ languages including templates, memory storage, and compiler/linker Experience with or knowledge...Relocation packageShift work- ...photonics, microfluidics, embedded systems, and semiconductor-grade... ...an Embedded AI Systems Engineer to architect and own the firmware... ...photonics, fluidics, mechanical, and software. What You’ll Do Build... ...optimize timing, control loops, memory, and performance What We’re...Summer workCasual work
$116.36k - $155.15k
...businesses connect, secure, and scale in an AI-driven world. By... ...independent efforts to all aspects of system integration including design,... ...in system architecture and engineering disciplines. Specific... ...peripherals, settings, directories, storage, etc. in accordance with...Full timeTemporary workRemote work1 day per week- ...and forward‑thinking AI Engineer to join our AI... ...and optimize intelligent systems that power automation,... ...databases (e.g., PGVector) and LLM orchestration tools to support retrieval and memory in generative systems.... ...production‑grade AI systems at scale in cloud environments....H1b
$152k - $241.5k
...for highly motivated Senior Software Engineers to work on our GPU Fabric Networking... ..., implement, and maintain system software that enables... ...communication hardware and software for large‑scale computing platforms. Work... ...NVIDIA GPUs. Knowledge of memory coherence and consistency...- Cotality is seeking a Principal Software Development Engineer to lead the design and testing of complex software systems. This senior position requires developing scalable, high-quality... ...transform existing services into high-scale intelligent platforms. The ideal...
$184k - $287.5k
We are looking for a Senior Software Engineer to help build NeMo Platform, NVIDIA's product for developing, evaluating, deploying, and operating AI systems at scale. This role will focus on NeMo Evaluator, which helps teams understand whether changes to AI agents are making...$116.3k - $184.2k
...opportunities to work on revolutionary systems that impact people's lives... ...develop, integrate and test software for our end-user customers... ...teams, such as with Systems Engineering, Cloud & Application, Test... ...including templates, memory storage, and compiler/linker Experience...Relocation packageShift work$121.4k - $218.6k
...Do you enjoy solving large scale distributed content delivery... ...-density hardware and software infrastructure... ...and performance of our systems. You'll define key performance... ...Site Reliability Engineer, you will be... ...advanced AI utilities and LLM-assisted development paradigms...Work experience placementWork at office- ...fastest and most widely used software load balancer. Organizations... ...observability and security at any scale and in any environment.... ...We are seeking a Software and Systems Engineer to join the small, elite core... ...Operations : Own and evolve our large-scale, declarative...Full timeImmediate startRemote workWorldwideShift work
$89k - $134k
...individual to join AIS as a Infrastructure Engineer. Core Knowledge & Skills: Manages... ...needs of our client as a Cloud Systems Engineer. Project Summary AIS is supporting... ...of the cloud-based networks, software systems for scaling and reliable operations. Performs system...- ...experienced technologist to design and operate cloud-native, distributed systems at scale. You will own architecture across frontend, backend, and data layers, guiding multiple teams and mentoring senior engineers. The role emphasizes security, reliability, and alignment with...Remote job
- ...We build and deploy advanced AI systems and autonomous technologies for... ...role is built for a strong engineer with a versatile software development background and a real... ...visual understanding, and VLM/LLM-powered reasoning. Turn large-scale visual data collected by autonomous...
- Accenture is seeking a Senior AI Engineering Leader to architect and deliver enterprise AI platforms across multiple model providers and... ...(0-100%), and you will collaborate with clients to shape large-scale AI programs, delivering scalable, secure, and governance-aligned...
$103k - $155k
...on Relativity’s legal AI software to securely surface and... ...What We Do At Relativity, engineers do more than write code - they build systems that enable users to... ...insights from complex data at scale using cloud-native... ...performance systems that process large volumes of data. This...Remote work- ...secure APIs, and collaborating with cross-functional teams. Strong CI/CD and debugging skills are expected to deliver high-quality software in a fast-paced environment. The candidate should have experience with .NET Core, ASP.NET Core, Cosmos DB, Azure, SQL Server,...Remote job
- ...Compute team to design, build, and scale the next generation of bare-metal provisioning systems powering millions of servers worldwide. As a senior engineer, you will develop highly reliable and... ...If you are interested in building large‑scale distributed infrastructure...WorldwideFlexible hours
- ...of Job Function As a Software Engineer, you will be a core contributor... ..., support production systems, and collaborate daily... ...that matters. Principal Duties and Essential Responsibilities... ...and enterprise-scale data volumes.... ...capabilities — including LLM-powered features,...Local areaWorldwideShift work
$83.43k - $222.48k
...a time. Position Summary As a Senior Software Engineer, you'll be a key member of a collaborative... ...of business‑critical, distributed systems. We’re looking for technically strong,... ...stack development to ship and operate large‑scale systems. Strong SQL skills and understanding...Hourly payFull timeTemporary workWork experience placementLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer - Large-Scale LLM Memory and Storage Systems. Be the first to apply!
- principal software engineer Oklahoma City, OK
- senior principal cloud computing engineer Oklahoma City, OK
- principal cloud computing engineer Oklahoma City, OK
- senior principal scientist Oklahoma City, OK
- principal Oklahoma City, OK
- software trainer Oklahoma City, OK
- ultimate software Oklahoma City, OK
- software sales executive Oklahoma City, OK
- software applications developer Oklahoma City, OK
- remote software sales Oklahoma City, OK


