Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Storage Software Engineer - DGX Cloud

$224k - $356.5k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.NVIDIA DGXC Storage team handles some of the fastest training and inference tasks. Every GPU cycle depends on a storage platform built to keep tens of thousands of accelerators continuously busy. It maintains exabytes of data securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With the growth of accelerated computing, storage is essential. It can make the difference between effective GPU use and wasted potential, and between launching a frontier model on time or missing the deadline by months. We’re looking for a hands-on Storage Software Engineer to join the storage team as an individual contributor and technical lead. You will contribute to open-source parallel and distributed file systems and keep our largest GPU clusters fast, reliable, and durable. You will stay deeply hands-on: writing and reviewing production code, chasing root causes in the field, and setting the configuration and tuning standards our GPU fleets run on. This is a chance to do foundational storage engineering for the AI era at the company that introduced accelerated computing.What you’ll be doing:Contribute to open-source file systems. Contribute code to open-source parallel and distributed file systems, and distributed object storage. Upstream fixes and features, and engage directly with the upstream communities and maintainers.Serve as a hands-on storage software lead. Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it. Make the final technical calls on storage deliveries against measurable targets.Triage and troubleshoot at scale. Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery.Validate architecture and capabilities. Validate storage architecture, capabilities, performance, and durability. Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets.Recommend configuration, tuning, and guidelines. Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them.Partner broadly. Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture.Work AI-first. Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations.What we need to see:BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience.Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale.Contributions to open-source projects involving a distributed or parallel file system.You are fully engaged in engineering tasks. You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others.Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI).Strong written and verbal communication; capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers.Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build.100% hands-on engineering. You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating.Ways to stand out from the crowd:Maintainers or sustained contributions to widely used public projects.Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks.Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience.Kubernetes and CSI driver development for storage.Hands-on experience with SPDK, libfabric, or FUSE performance optimization.NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. Our invention, the GPU, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is seeking exceptional individuals like you to help us drive the next wave of artificial intelligence. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Storage Software Engineer - DGX Cloud in Santa Clara, CA vacancy
  • $176k - $276k

    Production engineering is a field that involves crafting, building, and maintaining...  ...various areas, including software and systems engineering practices, storage, data management, and services. Professionals...  ..., along with open-source cloud-enabling technologies such as... 
    Senior
    Software
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...impact on the world.Are you passionate about building world-class reliability systems? Join NVIDIA as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a pivotal part of a team that redefines operational excellence. Our team is at the forefront of... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are looking for a Senior Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’s high-performance GPU infrastructure. You will play a critical role in designing scalable automation solutions, integrating diverse systems, and... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing the performance, efficiency, and resiliency...  ...end-to-end behavior across GPUs, networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads, establish... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI research...  ...workloads. We are looking for Senior Software Engineers to help build the automation, tooling...  ...follow-up work.Partner with platform, storage, networking, security, and workload teams... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    NVIDIA is hiring experienced software engineers to help scale up its AI Infrastructure. We expect you to have significant software engineering...  ...deep learning.What you will be doing:You will be part of an DGX Cloud team responsible for production systems that enable large... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is hiring experienced software engineers to help scale up its AI Infrastructure. We expect you to have significant software engineering...  ...deep learning.What you will be doing:You will be part of an DGX Cloud team responsible for production systems that enable large... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $168k - $264.5k

    NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient network to support NVIDIA software development workflows and tools, including CI/CD pipelines, compute resource management... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $224k - $356.5k

    NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation and Trust Services team builds the secure cloud services that show customers their NVIDIA platforms are healthy, resilient, and ready... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $136k - $224.25k

    NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence... 
    Senior
    Software
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

    NVIDIA's Object Storage Platform team builds and operates...  ...researchers and engineers to reliably store massive...  ...You will own the full software development and service...  ...engineering organization to senior leadership, providing...  ...platforms, or large-scale cloud data services; hands-on... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global...  ...service enablement. We build software and automation to...  ...are looking for a hands-on senior engineer to own the lifecycle and automation...  ...health, cluster networking, storage, scheduling, workload placement... 
    Senior
    Software
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that...  .... We are seeking an AI infrastructure software engineer to join our team. You'll be...  ...efficiency and availability of AI systems.As a senior DGX Cloud AI Infrastructure software engineer... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...technology—and amazing people.We are looking for a Principal Software Engineer to join our DGX Cloud team and build the foundational systems that drive...  ...-multiplier by coaching, mentoring, and encouraging senior engineers, elevating the technical standards and guidelines... 
    Software
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $200k - $322k

     ...program manager for NVIDIA's DGX Cloud. We want enthusiastic, diligent...  ...service providers and NVIDIA engineering teams, building outstanding...  ...execution of large programs, software engineering projects in a matrix...  ...cross org alignment across senior and executive leaders.Possess... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $200k - $322k

     ...NVIDIA’s DGX Cloud team helps some of the most advanced AI builders in the world move from...  ...recommendations across compute, networking, storage, and cloud environments.Turn repeat...  ...and reduce friction.gainsightWork across Engineering, Product, Operations, and Finance to surface... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $258.75k

     ...computing to deliver world-class technology. The DGX Cloud organization plays a pivotal role in this mission, crafting the software operating layer for NVIDIA's AI factory....  ...and incorporate findings into product and engineering plans.Own lab operations and partner... 
    Senior
    Software
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $200k - $322k

    NVIDIA is seeking a Senior Technical Program Manager to lead Trust Services programs for DGX Cloud. DGX Cloud powers large-scale AI infrastructure...  ...security, compliance, engineering execution, and partner...  ...across firmware, platform, and software teams.Establish program... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $258.75k

    The DGX Cloud TPM Team is seeking a Senior Technical Program Manager (TPM) to lead complex, multi-functional...  ...driving full-stack initiatives involving software, hardware, and cloud platforms. You...  ...NVIDIA’s top AI researchers and engineers in building and scaling new... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $184k - $287.5k

    At NVIDIA, the DGX Cloud division merges fresh hardware and software innovations to offer leading accelerated computing...  ...workloads worldwide. Our team of skilled engineers is committed to addressing major...  ...the world!We are looking for a Senior Systems Software Engineer with... 
    Senior
    Software
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We are looking for a Senior Software Engineer to become part of our storage management plane team. The management plane is a web-based application crafted to provide our storage customers the capabilities to handle and supervise our distributed storage infrastructure.... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $200k - $322k

    NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product...  .... The company seeks a Senior Technical Program Manager (TPM...  ...NVIDIA’s next-generation AI software platforms. In this role, you...  ...responsible for managing high-impact engineering programs within a dynamic,... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...impact on the world. The DGX Cloud organization at NVIDIA...  ...‑edge hardware and software innovation to deliver...  ...group of forward‑thinking engineers tackling some of the...  ...We’re searching for a Senior Systems Software...  ...architecture, networking, storage systems, and accelerator... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...NVIDIA DGX Cloud is scaling GPU infrastructure across internal, partner...  ...are looking for Principal Software Engineers to help shape the technical...  ...clusters.This role is for senior technical leaders who can define...  ...platform, infrastructure, storage, networking, security, and... 
    Software
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $208k - $327.75k

     ...seeking a world-class Senior Product Manager to...  ...While the NVIDIA DGX is the undisputed...  ...as the public cloud? The mission is to...  ...this role, own the software-defined blueprint...  ...InfiniBand/Ethernet), storage architectures, and...  ...of multiple engineering fields. As you define... 
    Senior
    Software
    Full time
    Night shift

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $272k - $431.25k

     ...NVIDIA DGX Cloud is the AI supercomputing-as-a-service substrate designed...  .... As a Security Data Engineer within our Infrastructure Security...  ...: Architect and run the storage layer. A data lake/lakehouse...  ...Production-Grade Coding: A strong software engineering background with... 
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $264.5k

    NVIDIA is seeking a Sr Network Security Engineer to implement and maintain robust security across on-premise and cloud environments - enabling business verticals that span Graphics Drivers to AI and Deep Learning. In this role, you will lead the deployment and management... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $183k - $247.6k

     ...words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment...  ...help. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security... 
    Senior
    Software
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ..., and delivery of enterprise storage and data protection platforms...  ...solutions across datacenter and cloud environments, while ensuring...  ...AI-scale workloads), and engineering team leadership.THE PERSON:Strong...  ..., etc.)Experience with open, software-defined, and vendor-neutral... 
    Senior
    Software

    AMD

    San Jose, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Storage Software Engineer - DGX Cloud. Be the first to apply!