Distinguished Engineer, Storage - AI Cloud
$320kNVIDIA
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.AI Cloud Data StorageNVIDIA DGXC Storage org handles some of the fastest training and inference tasks. Every GPU cycle depends on a storage platform built to keep tens of thousands of accelerators continuously busy. It maintains exabytes of data securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With the growth of accelerated computing, storage is essential. It can make the difference between effective GPU use and wasted potential and between launching a frontier model on time or missing the deadline by months. We seek a Distinguished Engineer to lead NVIDIA's storage strategy for AI Cloud across the Neocloud Provider (NCP) and Cloud Service Provider (CSP) ecosystem. You will direct the architecture of high-performance parallel file systems, object stores, and block storage at exabyte scale. You will stay hands-on, collaborating with engineers, SREs, partners, and storage vendors. You will apply NVIDIA's AI tools to increase your productivity and that of those you impact. This is a distinctive prospect to establish the storage framework of the AI era at the company that introduced accelerated computing.What you'll be doing:Lead the multi-year technical plan for AI Cloud Storage expansion across NCPs — determine the reference architecture, capabilities, performance and durability SLOs, qualification methodology, and roadmap for the high-performance file, object, and block storage that each NCP must offer to qualify for NVIDIA GPU allocation.Serve as the chief storage architect with deep hands-on involvement. Lead key reviews of storage builds and investigate root causes of complex production problems. Develop prototype reference implementations to minimize risks in new initiatives. Make final technical decisions on NCP storage deliveries using measurable SLOs. Apply AI tools heavily to amplify your technical influence throughout the program.Define the standard for "production-ready" in NCP storage, including durability and availability SLOs measured in 9s. Ensure sustained efficiency per TiB, observability, blast-radius containment, and reduced operational toil. Influence GPU delivery gating by requiring AI Cloud to accept GPU capacity only after verifying storage-focused ancillary services.Develop and guide the architectural direction by working closely with collaborators in training, inference, and accelerated-computing product lines. Coordinate with site-reliability, operations, networking, and security colleagues. Work together with external cloud providers, neocloud operators, and storage vendors to align on a common architecture.Develop the open-source path forward for AI storage. Establish and guide an open-source strategy that broadens the AI storage ecosystem. Advocate for a GitHub-first, security-first stance. Engage deeply with upstream open-source communities. Formalize the APIs, SDKs, and protocols allowing partners and the industry to build, integrate, and create with NVIDIA at the AI storage level.Lead an engineering culture centered on AI tools. Regularly use modern AI coding and agentic tools in your daily tasks. Show what 10 engineering means at NVIDIA. Distribute patterns, prompts, and evaluation harnesses across the storage organization.Partner with peer Distinguished and Principal storage architects across the organization to tackle the most difficult, long-term technical challenges. Make automation the only acceptable solution for infrastructure management tasks like live software upgrades, node and drive replacements, capacity rebalancing, cross-DC data movement, and dataset lifecycle. Establish root-cause analysis and corrective action rigor on every major incident. Design the storage layer for workloads spanning the next several GPU generations, including disaggregated inference with storage-backed KV caching, large-scale write-once-read-many inference patterns, exabyte regional object stores, and cross-DC dataset versioning and copy management.Mentor and develop senior, principal, and distinguished engineers across the storage organization and nearby business units. Raise the technical bar broadly. Represent NVIDIA externally in standards bodies, open-source communities, customer briefings, and industry forums (FAST, SC, OCP, SNIA, Linux Storage Summit).What we need to see:BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience.A minimum of 18+ years of practical engineering experience in storage technology is needed. This involves extensive involvement with a high-performance parallel file system like Lustre, GPFS / Spectrum Scale, WEKA, VAST, BeeGFS, DAOS, or its equivalent, handling data at multi-petabyte scale. Candidates must also have wide-ranging expertise in object storage (S3 / Swift-class) and block storage (NVMe-oF, NVMesh-class, iSCSI).A track record of crafting and managing storage platforms at exabyte scale for performance-critical workloads — AI training, HPC, video, or hyperscale data lakes — including direct responsibility for durability, availability, and performance SLOs measured in 9s.Demonstrated ability to set technical strategy across business units and partner organizations. You have driven multi-year storage architectures adopted by multiple teams, vendors, or customers. You can point to measurable outcomes such as GPU utility lift, $/PB reduction, incidents eliminated, and time-to-bring-up compressed.You are 100% hands-on in engineering. You write and review production code yourself. When a bug requires it, you read Lustre, NFS, kernel, NVMe-oF, or SPDK source code. You also run scale tests or recovery drills personally instead of delegating.Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).Frequent daily use of advanced AI coding and autonomous tools, including specific examples showing how you accelerated building, coding, debugging, validation, and operations. Also, share your perspective on future trends.Excellent written and verbal communication. You can write a one-pager that aligns a VP. You can also write a six-pager that aligns an entire org. You can explain a deep technical trade-off to an SRE, a vendor CTO, and an internal customer in the same week.Comfort operating in a 24/7 production environment where storage incidents directly impact GPU revenue, with a security-first approach baked into every build.Ways to stand out from the crowd:Proven background in designing or managing storage solutions for AI training or inference at 10k+ GPU scale, demonstrating clear improvements in GPU utilization or reducing I/O bottlenecks.Open-source contributions or maintainership in Lustre, NFS, SPDK, NVMe / NVMe-oF, CSI, Ceph, MinIO, RocksDB, or related projects.Built or led a disaggregated-inference or Inference-Time-Compute storage architecture — KV caching to fast in-cluster or GPU-adjacent storage, WORM at scale, storage-aware scheduling, or database-integrated inference.Public technical contributions — patents, peer-reviewed papers (FAST, SOSP, NSDI, OSDI, ATC), keynote talks, or RFCs — that demonstrate expertise and leadership in storage for AI infrastructure.NVIDIA led the way in accelerated computing. Today, our AI infrastructure drives global intelligence, changing industries worldwide. The AI Cloud Storage group forms the base that maintains the world's largest GPU fleet's productivity. Every model trained, every inference served, and every checkpoint saved passes through systems we develop, construct, and manage.Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$285k - $355k
...infrastructure—delivering unified storage, integrated data services, and solutions... ...full potential of their data, from AI to multicloud. Ready to innovate... ...Summary We are looking for a Distinguished Engineer (DE) reporting into Cloud Business Unit SVP&GM. Someone in...CloudPart timeWork at officeLocal area- Senior Distinguished Technologist - Pre-Sales AI & Data Center Networking This role has been designated as ‘Remote... ...Enterprise is the global edge-to-cloud company advancing the way people live... ...waves of innovation. Our Sales Engineering team empowers customers and partners...CloudFull timeWork experience placementLocal areaImmediate startRemote workWork from home
- Distinguished Technologist, Private Cloud AI - Applied & Agentic AIThis role has been designed as 'Hybrid' with a requirement that you will work on average... ...reusable AI components and agents and partner with engineering to take POCs into scalable, production‑grade services...CloudFull timeWork experience placementWork at officeLocal areaImmediate start2 days per week
$320k
NVIDIA is looking for a Distinguished Engineer to act as a senior technical leader in the Production... ...about cluster operations involving DGX Cloud GPU capacity.At NVIDIA, Production Engineering... ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes....CloudFull timeRemote work$169k - $338k
...Segment: Home OfficePosition Summary...What you'll do...As a Distinguished AI/ML Engineer within Walmart Global Tech’s Reliability Engineering... ...fault-tolerant systems and services across Walmart’s hybrid cloud infrastructure with emphasis on autonomous recovery and intelligent...CloudFull timeTemporary workPart time$168k - $270.25k
...at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing,... ...solutions while harnessing the power of cloud computing. You will be responsible for... ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes....CloudFull time$117k - $234k
...OfficePosition Summary...The Enterprise Storage Platform team, part of Walmart Global Tech’s Enterprise and Cloud organization, designs, engineers, and operates the next-generation storage... ...as Code, observability, and AI/AIOps-driven operational capabilities to...CloudFull timeTemporary workPart time$137k - $263k
...autonomous network vision.Key ResponsibilitiesDevice-Side AI & Telemetry: Designing, engineering, and deploying edge-side telemetry pipelines and AI/ML... ...cause analysis (RCA), and automated decision-making.Multi-Cloud & Data Architecture: Building and scaling data analytics...CloudFull timeTemporary workPart timeWork experience placementWork at officeWork from homeShift work3 days per week$267k - $356k
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from... ...work from home day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-class storage offerings...CloudWork experience placementWork at officeLocal areaWork from homeFlexible hours$320k
...tapping into the unlimited potential of AI to define the next era of... ...a lasting impact on the world.AI Cloud Data StorageNVIDIA DGXC Storage org handles some of the fastest training... ...the deadline by months. We seek a Distinguished Engineer to lead NVIDIA's storage strategy...CloudFull timeWorldwide$320k
...rapidly growing enterprise and cloud provider businesses. These... ...and a fully optimized NVIDIA AI and HPC software stack. We’re... ...Computer Science, Electrical Engineering or related field (or equivalent... ...DOCA)Knowledge of enterprise storage architectures and distributed...CloudFull timeShift work$176k - $276k
Production engineering is a field that involves crafting, building, and... ...systems engineering practices, storage, data management, and services... ...deployment, along with open-source cloud-enabling technologies such as... ...data access for HPC and AI/ML workloads.Storage Production...CloudFull timeFlexible hours$320k
...data centers is the ability to engineer integrated system designs in... ...a Global Connectivity Distinguished Engineer to accelerate next-generation... ...and subsea—that interconnect AI Factories. You will act as... ...leadership role within a Hyperscale Cloud Provider or a Tier-1 Global...CloudFull time- Distinguished Technologist, Distributed Systems & Agentic... ...is the global edge-to-cloud company advancing the... ...impact systems, mentors engineering teams, and helps define... ...and tools, including AI‑assisted coding workflows... ...(OLAP, distributed storage systems, open data formats...CloudFull timeWork experience placementWork at officeLocal areaImmediate start2 days per week
$320k
NVIDIA is seeking a Distinguished Engineer to serve as a senior technical leader in the Production Engineering... ...leading cluster activities within DGX Cloud GPU capacity. Production Engineering at... ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes....CloudFull timeLocal area$221.2k - $387.1k
...It all started when engineer Fred Luddy wrote code that automated... ...work. Today, ServiceNow is the AI control tower for business reinvention... ...will lead the Data & Storage Reliability Engineering organization... ...services, storage platforms, cloud infrastructure, and...CloudFull timeTemporary workWork at officeImmediate startRemote workFlexible hoursShift work- ...the world’s biggest companies and public cloud, Western Digital is fueling a brighter,... ...’ll find Western Digital supporting the storage infrastructure behind many of these platforms... ...tools.QualificationsBS or MS Degree in Engineering, Physics, Materials Science or a related...CloudTemporary workWork experience placementImmediate startRemote workFlexible hoursShift work
$57.69 - $96.15 per hour
..., United States / Toronto, CanadaProducts - Engineering /Fulltime /HybridOver 50,000 customers globally trust our end-to-end, cloud-driven networking solutions. They rely on our... ...Networks, Inc. (EXTR) is a global leader in AI-powered cloud networking, committed to...CloudFull timeH1bWork at officeRemote workWorldwideWork visa$224k - $356.5k
...tapping into the unlimited potential of AI to define the next era of computing. An era... ...Architect in the Agent Harness & Runtime Engineering team to build foundational systems for... ...distributed execution, data/ETL pipelines, HPC, cloud, Kubernetes, and GPU compute environments...CloudFull timeRemote work- ...the quantum revolution and the AI era. Join our team of creators... ...for Hardware Development Engineers to develop, test and provide customer... ..., IBM Power Systems, IBM Storage, and IBM Quantum Systems. Development... ...business and optimized for cloud computing. YOUR LIFE @ IBM...CloudFull timeContract workPart timeFixed term contractInternshipShift work
- Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This... ...10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude... ...skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge...Cloud
$320k
...into the unlimited potential of AI to define the next era of... ...rapidly growing enterprise and cloud provider businesses. Each bringing... ...technical leader to drive the engineering roadmap and innovation for our... ...architectures. Knowledge in storage and networking technologies.We...CloudFull timeShift work- ...NetApp seeks a cloud-performance engineer to design, develop, analyze, and optimize performance for cloud storage services and distributed systems. You will collaborate with cross-functional... ..., analytical, and creative, with hands-on AI/ML in performance engineering, strong OS...Cloud
$160k - $185k
...Supermicro is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big Data,... ...community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us....CloudWorldwide$150k - $230k
...industry leader in data-driven, client-to-cloud networking for large data center, campus... ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation... ..., LRO, and LPO—the next generation of AI data center networking solutions.Lead...Cloud$320k
...industry in delivering accelerated computing in cloud and enterprise environments. We’re a team of innovative engineers dedicated to solving some of the world’s biggest... ...development of our global strategy for scaled-out AI inferencing. You will architect the high-...CloudFull timeWorldwide$184k - $287.5k
...into the unlimited potential of AI to define the next era of... ...member of the HW Infrastructure Storage Strategy team, you will provide... ...requirements of an expanding cloud infrastructure. As an expert,... ...Computer Science, Electrical Engineering or related field or equivalent...CloudFull time$150k - $185k
...Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/... ..., passionate, and committed engineers, technologists, and business leaders... ....Job Summary:The world’s largest AI and cloud platforms are being powered...CloudWorldwide- ...Oracle Cloud Infrastructure’s Object Storage Service team is seeking a senior engineer to own software design and development for major components in a large-scale, distributed storage platform. You will be a hands-on coder who values simplicity, scalability, and collaborative...Cloud
$183k - $247.6k
...from foundational services such as Amazon’s Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations... ...change the world.We are seeking a Hardware Design Engineer with role in the definition, design and validation...CloudLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Distinguished Engineer, Storage - AI Cloud. Be the first to apply!
- senior cloud solutions architect Santa Clara, CA
- cloud engineer Santa Clara, CA
- senior principal cloud computing engineer Santa Clara, CA
- aws cloud architect Santa Clara, CA
- software engineer - cloud services Santa Clara, CA
- cloud engineering manager Santa Clara, CA
- senior cloud network engineer Santa Clara, CA
- principal cloud computing engineer Santa Clara, CA
- senior devops cloud engineer Santa Clara, CA
- aws cloud security engineer Santa Clara, CA





