Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

$272k - $431.25k

NVIDIA

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for fleet-scale reliability, working directly with engineering teams of key CSP / hyperscale customers to ensure NVIDIA platforms achieve target MTBI (Mean Time Between Interruptions) in production. In this role, you will augment NVIDIA's internal software/firmware and quality teams with a dedicated CSP-facing focus. You will drive work streams with CSP engineering teams to build shared understanding of reliability software/firmware architecture, methodology, incorporate their fleet telemetry and failure data into NVIDIA's improvement priorities, and validate that reliability improvements measured in the lab translate to real customer environments. Your cross-CSP visibility enables you to distinguish systemic architectural gaps from environmental or configuration-specific issues that no single customer engagement could identify alone.What you'll be doing:Drive reliability work streams with CSP engineering teams — ensuring shared understanding of MTBI measurement methodology, failure classification, and health monitoring architectureGather and synthesize CSP fleet reliability data — identify failure patterns that appear across multiple customers and champion improvements back into NVIDIA's firmware, driver, and hardware teamsDefine consistent MTBI measurement methodology that works across different CSP monitoring environments and operational practicesConduct fleet-scale failure pattern analysis using statistical methods (Pareto, survival analysis, Weibull) to classify failures as systemic, environmental, or configuration-specificDrive fleet health monitoring integration architecture — ensure NVIDIA's health agents, telemetry, and reporting align with CSP operational workflows and automationDefine burn-in reliability test environment and cluster certification criteria in collaboration with quality teams, validating with customers that criteria are meaningfulCollaborate with CSPs to ensure reliability-related integration work (health monitoring deployment, telemetry pipeline, alerting configuration) is complete ahead of at-scale launchDevelop predictive failure models using fleet telemetry and validate their effectiveness in customer environmentsWhat we need to see:15+ years of experience in systems software at datacenter scale, or reliability engineering with focus on at-scale challenges.BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience)Deep expertise in multi-NUMA, rack-scale system software and firmware. Statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classificationExperience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlationUnderstanding of hardware failure modes in large-scale GPU/accelerator deployments — ability to classify and prioritize across compute, interconnect, memory, power, and thermal domainsExperience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems. Familiarity with predictive maintenance or anomaly detection approaches applied to fleet health dataCustomer obsession — genuine passion for understanding fleet reliability challenges at scale and translating them into actionable engineering prioritiesStrong communication — ability to present statistical reliability findings to both deep technical audiences and executive leadership. Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams without direct authorityWays to stand out from the crowd:Experience in fleet reliability at a hyperscaler (hardware health, fleet reliability at leading CSP/Hyperscaler)Familiarity with NVIDIA GPU error taxonomy (Xid errors, NVLink error counters, thermal events, CPER records)Experience building health scoring or predictive failure models for accelerator or HPC infrastructureBackground in defining MTBI/MTBF measurement standards or certification programs for complex multi-component systemsUnderstanding of how reliability data flows from device firmware through telemetry pipelines to fleet-level dashboards and automated remediationNVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, hardworking and self-motivated, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements in Santa Clara, CA vacancy
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal...  ...to ensure they can reliably manage, update, and operate...  ...GPU firmware at fleet scale. You will drive work streams...  ...in Artificial Intelligence, High-Performance Computing... 
    Intelligence
    Fleet
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to...  ...monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate... 
    Fleet
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...'re looking for a Principal Engineer to join our CSP Engagements team as the technical...  ...NVIDIA rack-scale systems, GPU architectures...  ...teams have reliable baseline measurements...  ...configuration, software, or workload differences...  ...in Artificial Intelligence, High-Performance... 
    Intelligence
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...developments in Artificial Intelligence, High-Performance...  ...motivated Senior Software Engineers to join our Fabric...  ...on NVLink Rack-Scale Systems Stability & Reliability. In this role, you...  ...testing, and fleet support.Lead reliability...  ..., and Customer engagement teams to improve... 
    Intelligence
    Fleet
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $272k - $431.25k

     ...seeking a strategic and technically proficient Principal Software Engineer to join the Data Center Systems and Software CSP engagements team. As a leader and technologist, you...  ...groundbreaking developments in Artificial Intelligence, High-Performance Computing and... 
    Intelligence
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build...  ...development of software systems. These systems...  ...at rack and fleet scale. Build open...  ...internal deployments and CSP environments....  ...needs. Establish reliability, security, validation... 
    Fleet
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...0 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native...  ...track record debugging large-scale, cloud-native stacks across...  ...groundbreaking developments in Artificial Intelligence, High-Performance Computing... 
    Intelligence
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...is seeking a Senior Systems Software test (lead) Engineer to join our Cloud Service Provider (CSP) Engagements team, focusing on ML software...  ...from cluster to rack scale full-stack validation with customer...  ...developments in Artificial Intelligence, High-Performance Computing... 
    Intelligence
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $258.75k

     ...Manager to join the CSP Engagements team, focused on...  ...and embedded software leaders—including software engineering managers, technical...  ...successful large‑scale deployment of NVIDIA...  ..., observability, reliability, and hardware/...  ...developments in Artificial Intelligence, High-Performance... 
    Intelligence
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...is seeking a Senior Firmware Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such as...  ...performance optimization for large-scale data center environments....  ...groundbreaking developments in Artificial Intelligence, High-Performance Computing... 
    Intelligence
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter...  ...optimization for large-scale data center environments.Collaborate...  ...groundbreaking developments in Artificial Intelligence, High-Performance Computing and... 
    Intelligence
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...ROLE: We are seeking a Principal Software Engineer to serve as the...  ...stress, stability, scale-out, and system-level...  ...Jenkins / internal CI fleets, hardware lab orchestration...  ...strategic customer engagements, OEM qualification...  ...may use Artificial Intelligence to help screen,... 
    Intelligence
    Fleet
    Contract work
    Shift work

    AMD

    San Jose, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA DGX Cloud is scaling GPU infrastructure across...  ...We are looking for Principal Software Engineers to help shape the...  ..., automation, and reliability across large-scale GPU...  ..., or multi-cloud fleet operations.Experience...  ...developments in Artificial Intelligence, High-Performance... 
    Intelligence
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $101k - $161k

     ...computing, artificial intelligence, and software-defined networking...  ...awards, such as Best Engineering Team, Best Company...  ...re looking for Site Reliability Engineers to join our...  ...production systems at scale. We are responsible...  ...CloudVision service fleet, ensuring... 
    Intelligence
    Fleet

    Arista Networks

    Santa Clara, CA
    7 hours ago
  • $136k - $224.25k

     ...for a Senior Network Reliability Engineer to support and...  ...needs across the whole software stack for NVIDIA,...  ...Vehicles and Artificial Intelligence.In this role, the...  ...be responsible for engaging with external...  ...resolution in large-scale networks and CSP environments, outstanding... 
    Intelligence
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $320k

     ...Distinguished Engineer to lead...  ...Provider (CSP) ecosystem...  ...exabyte scale. You will...  ...with site-reliability, operations...  ...first stance. Engage deeply...  ...Distinguished and Principal storage...  ...like live software upgrades,...  ...global intelligence, changing...  ...largest GPU fleet's... 
    Intelligence
    Fleet
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...developments in Artificial Intelligence, High-Performance Computing...  ...Experience supporting large‑scale HPC clusters using Slurm,...  ...host lifecycle management, fleet reliability/auto-healing, E2E observability...  ..., or Ruby.Mentored other engineers and influenced technical... 
    Intelligence
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...Principal Systems Software Engineer at NVIDIA is an engineering discipline...  ...and maintain large scale production systems...  ...services run maximum reliability and uptime as...  ...logging and alerting Engage in and improve the...  ...infrastructure powers global intelligence, transforming every... 
    Intelligence

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $262k - $365k

     ...new features.Engage with partner teams...  ...of software solutions. Uphold...  ...innovation, growth, engineering excellence,...  ...information at massive scale, and extend...  ..., artificial intelligence, natural...  ...architecture of Google’s Fleet planning and...  ..., efficiency, reliability and velocity.... 
    Intelligence
    Fleet
    Temporary work
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $224k - $356.5k

     ...hands-on Storage Software Engineer to join the storage...  ...clusters fast, reliable, and durable. You...  ...standards our GPU fleets run on. This is a...  ...features, and engage directly with the...  ...troubleshoot at scale. Triage, troubleshoot...  ...in Artificial Intelligence, High-Performance... 
    Intelligence
    Fleet
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $2,000 per month

     ...for frontier intelligence. We co-design chips, racks, software, and manufacturing...  ...by leading engineers, Etched is redefining...  ...Product Reliability to lead...  ...architecture through fleet deployment....  ...reliability at scale. Key Responsibilities...  ...qualification engagement Background... 
    Intelligence
    Fleet
    Contract work
    Work at office
    Relocation package

    Etched

    San Jose, CA
    29 days ago
  • $142.8k - $274.8k

     ...than 25%Profession: Software EngineeringDiscipline...  ...Artificial Intelligence Cloud Inference team...  ..., and Dynamics.As a Principal Software Engineer - Performance on the...  ...APIs to enable large scale training and inferencing...  ...footprint of the computing fleet and achieve Azure AI... 
    Intelligence
    Fleet
    Ongoing contract
    Work at office
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    3 days ago
  • $272k - $431.25k

     ...team member to build planet-scale maps supporting self-...  ...work with a diverse team of engineers in mapping, perception, reconstruction...  ...pipelines that transform fleet data into reliable map products used in self-...  ...production-quality software systems.Solid foundation in... 
    Fleet
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $190k - $210k

     ...United StatesProducts - Engineering /Fulltime /HybridOver...  ...of position: Principal Software EngineerPosition type...  ...shape the future of intelligent networking, generative...  ...be scalable, secure, reliable, observable, adaptable...  ...architecture of large-scale distributed systems,... 
    Intelligence
    Full time
    H1b
    Local area
    Work from home
    Work visa
    Shift work

    Extreme Networks, Inc.

    San Jose, CA
    1 day ago
  • $114.6k - $234.6k

     ...Infrastructure (OCI) is looking for a Principal Software Engineer to lead the development of...  ...across OCI’s global fleet. As a Principal Engineer,...  ...to ensure OCI can launch, scale, and maintain new server...  ...operational overhead and high reliability. This role is ideal for... 
    Fleet
    Temporary work
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  •  ...are actively engaged in...  ...than great engineering—it takes a team...  ...security and reliability of the digital...  ...new kind of intelligence layer for AI...  ...to full rack scale. At the heart...  ...replaces reactive, software-bound...  ...are seeking a Principal Product Manager...  ...scales into fleet-level... 
    Intelligence
    Fleet
    Local area

    Axiado Corporation

    San Jose, CA
    7 hours ago
  •  ...driving a unified ROCm software stack across AMD’s...  ...NPI) at company-wide scale. The ideal candidate...  ...Workload Performance Engineering: Lead the profiling,...  ...reporting. Customer Engagement: Partner with top customers...  ...may use Artificial Intelligence to help screen, assess... 
    Intelligence

    AMD

    San Jose, CA
    7 hours ago
  • $208k - $333.5k

     ...network architecture & engineering. This is a hands-...  ...of global‑scale backbone and data...  ...that serve large fleets of CPU‑based compute...  ..., vendor engagement, and lifecycle strategy...  ...compliance, and reliability standards for all...  ...Engineering, Artificial Intelligence, Data Science,... 
    Intelligence
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $147k - $237.5k

     ...TeamEngineering - Our engineering team is at the...  ....Your CareerAs a Principal Engineer on the...  ...improvements in reliability.This is a rare opportunity...  ...highly scalable software features and...  ...configuration at scale for thousands of...  ...assistants, intelligent troubleshooting agents... 
    Intelligence
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...amazing people.We are looking for a Principal Software Engineer to join our DGX Cloud team and...  ...systems that automate fleet lifecycle operations at a massive scale.Drive technical alignment across...  ...integration, clear interfaces, and reliable end-to-end workflows, with a strong... 
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements. Be the first to apply!