Principal Software Engineer, Rack-Scale System Software — CSP Engagements (Santa Clara)
$272k - $431.25kNvidia
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams with dedicated CSP-facing technical leadership. Your focus is on the system-level software that manages, monitors, and recovers the rack as a whole — fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and SW-driven serviceability. You will drive work streams with CSP engineering teams to build shared understanding of the architecture, incorporate their operational feedback, and ensure integration readiness.What you'll be doing:Drive rack-scale SW/FW architecture alignment across CSP engagements — including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestrationDrive technical work streams with CSP engineering teams on rack-scale system software — ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operationCapture and synthesize CSP engineering feedback on rack-scale system software — health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior — champion that feedback into NVIDIA's architecture decisionsCollaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware developmentIdentify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices — drive documentation, tooling, and test strategy improvements as a resultCollaborate with execution teams on left-shift strategy — ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availabilityMake critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teamsWhat we need to see:15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience)Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability / reliabilityExperience with fabric management software, cluster management, or system-level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery)Understanding of error handling and recovery design patterns in distributed systems — fault isolation, retry policies, graceful degradationExperience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observabilityUnderstanding of GPU or accelerator system software (drivers, device management, power management) is a strong plusCustomer obsession — genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experienceProven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication — ability to translate complex system software architecture into actionable mentorship for customer engineering teamsWays to stand out from the crowd:Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management softwareBackground in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms)Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices)Familiarity with GPU or accelerator fleet operations — driver lifecycle, firmware rollout strategies, health-based schedulingUnderstanding of how system software decisions impact serviceability, availability, and operational cost at fleet scaleNVIDIA’s invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and establish teams with the most thoughtful people in the world. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement.NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Remote; US, CA, RemoteType: Full time
$272k - $431.25k
...'re looking for a Principal Software Engineer to join our CSP Engagements team as the technical... ...firmware and GPU system software, working... ...firmware at fleet scale. You will drive... ...multiple GPUs in a rack-scale systemUnderstanding... ...: US, CA, Santa Clara; US, TX, Austin; US...SuggestedFull timePart timeRemote work$272k - $431.25k
...'re looking for a Principal Software Engineer to join our CSP Engagements team as the technical... ...point for fleet-scale reliability,... ...you to distinguish systemic architectural gaps... ...expertise in multi-NUMA, rack-scale system... ...SummaryLocation: US, CA, Santa ClaraType: Full...SuggestedFull timePart time$272k - $431.25k
...the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems... ...internal deployments and CSP environments.Bridge... ...SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX,...SuggestedFull timePart timeRemote workShift work$272k - $431.25k
...'re looking for a Principal Engineer to join our CSP Engagements team as the technical... ...and drive systemic improvements in documentation... ...the latest NVIDIA rack-scale systems, GPU... ...identify configuration, software, or workload... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin;...SuggestedFull timePart timeRemote work$272k - $431.25k
...seeking a strategic and technically proficient Principal Software Engineer to join the Data Center Systems and Software CSP engagements team. As a leader and technologist, you... ...other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...SuggestedFull timePart timeShift work$272k - $431.25k
...many accelerators feel like a single system at datacenter scale. As large language models rapidly... ...LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap for... ...by law.SummaryLocation: US, CA, Santa Clara; US, WA, Remote; US, MA, RemoteType...Full timePart timeLocal areaRemote work$320k
...single node HGX/DGX systems all the way up... ...NVLink domain rack architectures.... ...AI and HPC software stack. We're searching... ...to drive the engineering roadmap and... ...internally and engage with industry leading... ...NVIDIA's rack-scale... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType...Full timePart timeShift work$272k - $431.25k
...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU accelerated data processing has moved from... ...characteristic protected by law.#deeplearningSummaryLocation: US, CA, Santa Clara; US, IL, ChampaignType: Full time...Full timePart timeWork experience placement$272k - $431.25k
...MODS organization seeks a Principal Engineer to architect and scale next-generation L10 and L11 diagnostic systems for Cloud Service Providers... ...systems and hardware / software interfaces is essential for... ....SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full...Full timePart time$248k - $391k
...seeking a highly skilled Principal Software Engineer to join our dynamic... ...-class AI inference systems. Join us in this... ...inference platform scaling to frontier-class models... ...for pre-release, rack-scale GPU systems (including... ...: US, CA, Santa Clara; US, RemoteType: Full...Full timePart time$272k - $431.25k
...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team... ...and managing workloads at production scale. Drive triage of the most difficult... ...characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart timeShift work$272k - $431.25k
...accelerated computing. As a Principal Software Engineer, you will lead the... ...of AI networking systems. You will apply your... ...complex customer engagements and help develop our... ...performance tuning at scale.Experience building... ...: US, CA, Santa Clara; US, CA, RemoteType:...Full timePart time$272k - $431.25k
...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations... ..., development, optimization, and scaling of foundational software technologies... ...characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$272k - $431.25k
...productivity required for strong scaling for HPC and generative AI... ...are looking for expert engineers to come and help design rack level solutions for next... ...complexity and project system resource requirements.... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...Full timePart time$208k - $260k
...educational organizations. As a Principal Software Engineer on the Network Management System team, you will lead the design... ...This role is based out of our Santa Clara, CA headquarters, following a hybrid... ...platforms that support large-scale deployments and long-term...Local areaWorldwide3 days per week$272k - $431.25k
...infrastructure at scale. Enterprise... ...Datacenter Software Tools team at... ...including system architects, firmware... ...customers of rack-scale... ...SWQA, Product engineering to left-shift... ...across different CSP environments.... ...and engage in collaborative... ...SummaryLocation: US, CA, Santa Clara; US, CA,...Full timePart timeRemote workShift workNight shift$272k - $431.25k
...service that runs in production, scales across cluster environments... ...hands-on delivery across system software, drivers, and CUDA to make... ...technical direction for an engineering team; mentor engineers,... ...law.SummaryLocation: US, CA, Santa Clara; US, TX, AustinType: Full...Full timePart time$272k - $431.25k
...member to build planet-scale maps supporting self-... ...fusion, and large-scale systems. You will transform... ...with a diverse team of engineers in mapping, perception... ...building production-quality software systems.Solid... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full...Full timePart timeWorldwide$272k - $431.25k
...people.We are looking for a Principal Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’s high-... ...lifecycle operations at a massive scale.Drive technical alignment... ....SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full...Full timePart time$272k - $431.25k
...working for us! The Cloud Engineering & Services team is... ...layer for agentic systems: signed policy, runtime... ...We are looking for a Principal Software Engineer, Agent... ...with Runtime Owners: Engage alongside OpenShell and... ...SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US,...Full timeContract workPart timeRemote work$120k - $275k
...solutions from silicon to systems including hardware and software to train and run the... ...micro-architects and design engineers to join our team as we... ...a Senior Platform TPM, Rack-Scale AI Systems to drive design... ...external CM/JDM partner engagement. You will work closely...Full timeContract workWork experience placementLocal areaRemote workMonday to FridayFlexible hours$272k - $431.25k
...seeking a highly skilled Principal Engineer to join our world-... ...scalability of the systems that power our global... ...deliver impact at global scale—all within an... ...regulated enterprise software delivery.Your base salary... ...SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType:...Full timePart time$272k - $431.25k
...and NVIDIA proprietary software, GeForce NOW... ...seeking an experienced Systems Software Engineering Manager to lead a team... ...for growing usage and scale of existing customers... ...concepts.Experience engaging with analysts, industry... ...SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart timeLocal area$272k - $431.25k
...Rubin–class compute platform engineered for low-Earth orbit mission.... ...architect to own end-to-end system software architecture for Space-1 and... ...platform software for large-scale data centers or mission-... ...law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...Full timePart timeWork experience placementRemote work$272k - $431.25k
...NVIDIA DGX Cloud is scaling GPU infrastructure across internal... .... We are looking for Principal Software Engineers to help shape the technical... ...influence, build critical systems, and turn ambiguous infrastructure... ....SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...Full timePart time$320k
...NVIDIA data center systems, such as DGX and... ...NVIDIA AI and HPC software stack. We’re looking... ...internally and engage with industry leading... ...in large-scale, collaborative environments... ..., Electrical Engineering or related field (... ...SummaryLocation: US, CA, Santa ClaraType: Full...Full timePart timeShift work- ...DescriptionGeneral information:Hitachi Energy is seeking for a Systems Integration Engineer for it's Santa Clara, CA or Houston, TX locations. This role is... ...environments.SummaryBuild and solve Linux production systems at scale. Multi-datacenter deployment and supporting a large-...Part time
$221.2k - $387.1k
...DescriptionIt all started when engineer Fred Luddy wrote code... ...DescriptionPrincipal Software EngineerThe... ...operations at enterprise scale. Together, these capabilities... ...do in this role: As a Principal Software Engineer, you... ...-scale distributed systems, data ingestion,...Part timeWork at officeImmediate startRemote workFlexible hours$272k - $431.25k
...personal assistants and engineering-productivity tools... .... Now we need a principal-level, hands-on... ...harden production systems and the... ...vision to ensure they scale.What you'll be doing... ...behave like mature software, not prototypes.Build... ...SummaryLocation: US, CA, Santa ClaraType: Full...Full timePart timeLive in$208k - $260k
...organizations. As a Principal Security Detections Engineer on the GigaSMART team,... ..., and high-performance systems software to identify suspicious activity... ...is based out of our Santa Clara, CA headquarters,... ...telemetry with strong focus on scale, accuracy, and...Local areaWorldwide3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer, Rack-Scale System Software — CSP Engagements (Santa Clara). Be the first to apply!
- principal software engineer Santa Clara, CA
- operating system engineer Santa Clara, CA
- computer system validation engineer Santa Clara, CA
- system performance engineer Santa Clara, CA
- system engineer contract Santa Clara, CA
- senior linux systems engineer Santa Clara, CA
- digital communications systems engineer Santa Clara, CA
- sr systems engineer Santa Clara, CA
- application system engineer Santa Clara, CA
- healthcare systems engineer Santa Clara, CA




