Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
$272k - $431.25kNVIDIA
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for fleet-scale reliability, working directly with engineering teams of key CSP / hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production. In this role, you will augment NVIDIA's internal software/firmware and quality teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of reliability software/firmware architecture, methodology, incorporate their fleet telemetry and failure data into NVIDIA's improvement priorities, and validate that reliability improvements measured in the lab translate to real customer environments. Your cross-CSP visibility enables you to distinguish systemic architectural gaps from environmental or configuration-specific issues that no single customer engagement could identify alone.What you'll be doing:Drive reliability work streams with CSP engineering teams — ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architectureGather and synthesize CSP fleet reliability data — identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teamsDefine consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practicesConduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specificDrive fleet health monitoring integration architecture — ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automationDefine burn-in reliability test environment and cluster certification criteria in collaboration with quality teams, validating with customers that criteria are meaningfulCollaborate with CSPs to ensure reliability-related integration work (health monitoring deployment, telemetry pipeline, alerting configuration) is complete ahead of at-scale launchDevelop predictive failure models using fleet telemetry and validate their effectiveness in customer environmentsWhat we need to see:15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges.BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience)Deep expertise in multi-NUMA, rack-scale system software and firmware. Statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classificationExperience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlationUnderstanding of hardware failure modes in large-scale GPU/accelerator deployments — ability to classify and prioritize across compute, interconnect, memory, power, and thermal domainsExperience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems. Familiarity with predictive maintenance or anomaly detection approaches applied to fleet health dataCustomer obsession — genuine passion for understanding fleet reliability challenges at scale and translating them into actionable engineering prioritiesStrong communication — ability to present statistical reliability findings to both deep technical audiences and executive leadership. Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authorityWays to stand out from the crowd:Experience in fleet reliability at a hyperscaler (hardware health, fleet reliability at leading CSP/Hyperscaler)Familiarity with NVIDIA GPU error taxonomy (Xid errors, NVLink error counters, thermal events, CPER records)Experience building health scoring or predictive failure models for accelerator or HPC infrastructureBackground in defining MTBI/MTBF measurement standards or certification programs for complex multi-component systemsUnderstanding of how reliability data flows from device firmware through telemetry pipelines to fleet-level dashboards and automated remediationNVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal... ...to ensure they can reliably manage, update, and operate... ...GPU firmware at fleet scale. You will drive work streams... ...in Artificial Intelligence, High-Performance Computing...IntelligenceFleetFull timeRemote work$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to... ...monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate...FleetFull timeRemote workShift work$272k - $431.25k
...'re looking for a Principal Engineer to join our CSP Engagements team as the technical... ...NVIDIA rack-scale systems, GPU architectures... ...teams have reliable baseline measurements... ...configuration, software, or workload differences... ...in Artificial Intelligence, High-Performance...IntelligenceFull timeRemote work$152k - $241.5k
...developments in Artificial Intelligence, High-Performance... ...motivated Senior Software Engineers to join our Fabric... ...on NVLink Rack-Scale Systems Stability & Reliability. In this role, you... ...testing, and fleet support.Lead reliability... ..., and Customer engagement teams to improve...IntelligenceFleetFull timeRemote work$272k - $431.25k
...seeking a strategic and technically proficient Principal Software Engineer to join the Data Center Systems and Software CSP engagements team. As a leader and technologist, you... ...groundbreaking developments in Artificial Intelligence, High-Performance Computing and...IntelligenceFull timeShift work$272k - $431.25k
...world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build... ...development of software systems. These systems... ...at rack and fleet scale. Build open... ...internal deployments and CSP environments.... ...needs. Establish reliability, security, validation...FleetFull timeRemote workShift work$184k - $287.5k
...0 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native... ...track record debugging large-scale, cloud-native stacks across... ...groundbreaking developments in Artificial Intelligence, High-Performance Computing...IntelligenceFull timeRemote work$184k - $287.5k
...is seeking a Senior Systems Software test (lead) Engineer to join our Cloud Service Provider (CSP) Engagements team, focusing on ML software... ...from cluster to rack scale full-stack validation with customer... ...developments in Artificial Intelligence, High-Performance Computing...IntelligenceFull timeLocal area$168k - $258.75k
...Manager to join the CSP Engagements team, focused on... ...and embedded software leaders—including software engineering managers, technical... ...successful large‑scale deployment of NVIDIA... ..., observability, reliability, and hardware/... ...developments in Artificial Intelligence, High-Performance...IntelligenceFull time$184k - $287.5k
...is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as... ...performance optimization for large-scale data center environments.... ...groundbreaking developments in Artificial Intelligence, High-Performance Computing...IntelligenceFull time$184k - $287.5k
NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter... ...optimization for large-scale data center environments.Collaborate... ...groundbreaking developments in Artificial Intelligence, High-Performance Computing and...IntelligenceFull time- ...ROLE: We are seeking a Principal Software Engineer to serve as the... ...stress, stability, scale-out, and system-level... ...Jenkins / internal CI fleets, hardware lab orchestration... ...strategic customer engagements, OEM qualification... ...may use Artificial Intelligence to help screen,...IntelligenceFleetContract workShift work
$272k - $431.25k
...NVIDIA DGX Cloud is scaling GPU infrastructure across... ...We are looking for Principal Software Engineers to help shape the... ..., automation, and reliability across large-scale GPU... ..., or multi-cloud fleet operations.Experience... ...developments in Artificial Intelligence, High-Performance...IntelligenceFleetFull time$101k - $161k
...computing, artificial intelligence, and software-defined networking... ...awards, such as Best Engineering Team, Best Company... ...re looking for Site Reliability Engineers to join our... ...production systems at scale. We are responsible... ...CloudVision service fleet, ensuring...IntelligenceFleet$136k - $224.25k
...for a Senior Network Reliability Engineer to support and... ...needs across the whole software stack for NVIDIA,... ...Vehicles and Artificial Intelligence.In this role, the... ...be responsible for engaging with external... ...resolution in large-scale networks and CSP environments, outstanding...IntelligenceFull timeRemote workShift work$320k
...Distinguished Engineer to lead... ...Provider (CSP) ecosystem... ...exabyte scale. You will... ...with site-reliability, operations... ...first stance. Engage deeply... ...Distinguished and Principal storage... ...like live software upgrades,... ...global intelligence, changing... ...largest GPU fleet's...IntelligenceFleetFull timeWorldwide$152k - $241.5k
...developments in Artificial Intelligence, High-Performance Computing... ...Experience supporting large‑scale HPC clusters using Slurm,... ...host lifecycle management, fleet reliability/auto-healing, E2E observability... ..., or Ruby.Mentored other engineers and influenced technical...IntelligenceFleetFull time$272k - $431.25k
...Principal Systems Software Engineer at NVIDIA is an engineering discipline... ...and maintain large scale production systems... ...services run maximum reliability and uptime as... ...logging and alerting Engage in and improve the... ...infrastructure powers global intelligence, transforming every...Intelligence$262k - $365k
...new features.Engage with partner teams... ...of software solutions. Uphold... ...innovation, growth, engineering excellence,... ...information at massive scale, and extend... ..., artificial intelligence, natural... ...architecture of Google’s Fleet planning and... ..., efficiency, reliability and velocity....IntelligenceFleetTemporary workWorldwide$224k - $356.5k
...hands-on Storage Software Engineer to join the storage... ...clusters fast, reliable, and durable. You... ...standards our GPU fleets run on. This is a... ...features, and engage directly with the... ...troubleshoot at scale. Triage, troubleshoot... ...in Artificial Intelligence, High-Performance...IntelligenceFleetFull timeWorldwide$2,000 per month
...for frontier intelligence. We co-design chips, racks, software, and manufacturing... ...by leading engineers, Etched is redefining... ...Product Reliability to lead... ...architecture through fleet deployment.... ...reliability at scale. Key Responsibilities... ...qualification engagement Background...IntelligenceFleetContract workWork at officeRelocation package$142.8k - $274.8k
...than 25%Profession: Software EngineeringDiscipline... ...Artificial Intelligence Cloud Inference team... ..., and Dynamics.As a Principal Software Engineer - Performance on the... ...APIs to enable large scale training and inferencing... ...footprint of the computing fleet and achieve Azure AI...IntelligenceFleetOngoing contractWork at officeLocal area3 days per week$272k - $431.25k
...team member to build planet-scale maps supporting self-... ...work with a diverse team of engineers in mapping, perception, reconstruction... ...pipelines that transform fleet data into reliable map products used in self-... ...production-quality software systems.Solid foundation in...FleetFull timeWorldwide$190k - $210k
...United StatesProducts - Engineering /Fulltime /HybridOver... ...of position: Principal Software EngineerPosition type... ...shape the future of intelligent networking, generative... ...be scalable, secure, reliable, observable, adaptable... ...architecture of large-scale distributed systems,...IntelligenceFull timeH1bLocal areaWork from homeWork visaShift work$114.6k - $234.6k
...Infrastructure (OCI) is looking for a Principal Software Engineer to lead the development of... ...across OCI’s global fleet. As a Principal Engineer,... ...to ensure OCI can launch, scale, and maintain new server... ...operational overhead and high reliability. This role is ideal for...FleetTemporary workFlexible hours- ...are actively engaged in... ...than great engineering—it takes a team... ...security and reliability of the digital... ...new kind of intelligence layer for AI... ...to full rack scale. At the heart... ...replaces reactive, software-bound... ...are seeking a Principal Product Manager... ...scales into fleet-level...IntelligenceFleetLocal area
- ...driving a unified ROCm software stack across AMD’s... ...NPI) at company-wide scale. The ideal candidate... ...Workload Performance Engineering: Lead the profiling,... ...reporting. Customer Engagement: Partner with top customers... ...may use Artificial Intelligence to help screen, assess...Intelligence
$208k - $333.5k
...network architecture & engineering. This is a hands-... ...of global‑scale backbone and data... ...that serve large fleets of CPU‑based compute... ..., vendor engagement, and lifecycle strategy... ...compliance, and reliability standards for all... ...Engineering, Artificial Intelligence, Data Science,...IntelligenceFleetFull time$147k - $237.5k
...TeamEngineering - Our engineering team is at the... ....Your CareerAs a Principal Engineer on the... ...improvements in reliability.This is a rare opportunity... ...highly scalable software features and... ...configuration at scale for thousands of... ...assistants, intelligent troubleshooting agents...IntelligenceFull timeWork at office$272k - $431.25k
...amazing people.We are looking for a Principal Software Engineer to join our DGX Cloud team and... ...systems that automate fleet lifecycle operations at a massive scale.Drive technical alignment across... ...integration, clear interfaces, and reliable end-to-end workflows, with a strong...FleetFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements. Be the first to apply!
- principal software engineer Santa Clara, CA
- principal Santa Clara, CA
- senior principal cloud computing engineer Santa Clara, CA
- principal architect Santa Clara, CA
- senior principal scientist Santa Clara, CA
- principal cloud computing engineer Santa Clara, CA
- software engineer - cloud services Santa Clara, CA
- id software Santa Clara, CA
- healthcare software sales Santa Clara, CA
- software technical support Santa Clara, CA

