Principal Software Engineer - Rack Scale Systems Infrastructure
$272k - $431.25kNVIDIA
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.What You Will Be Doing:Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.What We Need To See:BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.Ways To Stand Out from the crowd:Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US, Remote; US, MA, RemoteType: Full time
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role...SuggestedFull timeRemote workShift work$272k - $431.25k
...accelerators feel like a single system at datacenter scale. As large language models rapidly... ...edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap... ...storage, or ML systems infrastructure in C/C++ and Python, with a track...SuggestedFull timeLocal areaRemote work$226k - $369k
...part of our world-class software engineering team, you will take... ...the next-generation infrastructure and platforms for LinkedIn... ..., API design and systems design, and your... ...performs at massive scale. LinkedIn has pioneered... ...within our company.As a Principal Staff Software Engineer...SuggestedFor contractorsWork at officeFlexible hours$152k - $241.5k
...looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus on NVLink Rack-Scale Systems Stability & Reliability. In this role... ...diagnostics, recovery, and large-scale AI infrastructure, contributing directly to the...SuggestedFull timeRemote work$248k - $391k
...seeking a highly skilled Principal Software Engineer to join our dynamic... ...performance of our infrastructure both on-prem and in... ...-class AI inference systems. Join us in this... ...inference platform scaling to frontier-class models... ...for pre-release, rack-scale GPU systems (including...SuggestedFull time$226k - $369k
...approval. We are seeking a Principal Staff Software Engineer to join our organization.... ...-performing at global scale.As a Principal Staff Engineer... ...driving next-generation infrastructure that powers AI-first... ...large-scale infrastructure systems for end-to-end software lifecycle...For contractorsWork at officeRemote workWork from homeFlexible hours$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as... ...point for GPU firmware and GPU system software, working directly... ...NVIDIA GPU firmware at fleet scale. You will drive work streams... ...across multiple GPUs in a rack-scale systemUnderstanding of...Full timeRemote work$272k - $431.25k
...We're looking for a Principal Software Engineer to join our CSP Engagements team... ...technical focal point for fleet-scale reliability, working... ...enables you to distinguish systemic architectural gaps from environmental... ...expertise in multi-NUMA, rack-scale system software and...Full time- ...are at the forefront of software and hardware innovation,... ...days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat... ...You are able to build and scale software deliverables in... ...build out the deployment infrastructure, working closely with other...3 days per week
$174k - $253k
Write and test product or system development code.... ...code developed by other engineers and provide feedback... ...maintaining, or launching software products, and 1 year... ...at massive scale, and extend well beyond... ...built by the Technical Infrastructure team to keep it running...Worldwide$272k - $431.25k
...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU... ...processing problems challenges at large scale Provide recommendations and... ...decisions surrounding topics such as infrastructure, continuous integration and testing...Work experience placement$184k - $287.5k
We are looking for a senior systems software engineer to improve the operation and user experience of distributed system infrastructure using AI. We are passionate about the opportunity... ...AI-powered infrastructure that scales to run millions of requests and jobs on...Full time$249k
...for travelers everywhere.Principal Software Development Engineer — Platform & InfrastructureIntroduction... ...excellence of our cloud infrastructure and platform capabilities... ...who can design complex systems, implement solutions in... ...engineering for high-scale workloadsOwn and...Full time$152k - $241.5k
...Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous... ...of hardware-in-the-loop (HIL) infrastructure to support the robust deployment and... ...next-generation autonomous systems at scale.Build and evolve agentic AI tools and...Full time$272k - $431.25k
...Principal Systems Software Engineer at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the... ...experience15+ years of experience with Infrastructure automation, distributed systems...Full time$132k - $189k
Innovate and design in-rack and data center power... ...(PDU), battery backup system, etc. Analyze performance... ...degree in Electrical Engineering, Computer Engineering,... ...Engineer in Technical Infrastructure, you will play a key... ...builds the hardware and software technologies that...Worldwide$174k - $253k
Write and test product or system development code.... ...code developed by other engineers and provide feedback... ...maintaining, or launching software products, and 1 year... ...at massive scale, and extend well beyond... ...solutions.The AI and Infrastructure team is redefining what...Worldwide$147k - $211k
...code developed by other engineers and provide feedback... ...feedback. Triage product or system issues and debug/track... ...and resolution, and software test engineering.In... ...solutions.The AI and Infrastructure team is redefining what... ...at unparalleled scale, efficiency, reliability...Worldwide$184k
...world. We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely... ...performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies...Full time$114.6k - $234.6k
Oracle Cloud Infrastructure (OCI) is looking for a Principal Software Engineer to lead the development of scalable... ...secure infrastructure systems that underpin the core... ...server lifecycle from rack integration and hardware... ...ensure OCI can launch, scale, and maintain new server...Temporary workFlexible hours$272k - $431.25k
We're looking for a Principal Engineer to join our CSP Engagements... ...patterns and drive systemic improvements in... ...the latest NVIDIA rack-scale systems, GPU architectures... ...configuration, software, or workload differences... ...work (profiling infrastructure, benchmark harnesses...Full timeRemote work$272k - $431.25k
...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team builds the foundational software that turns... ..., monitoring, and managing workloads at production scale. Drive triage of the most difficult sequencing, initialization...Full timeShift work$258k - $387k
...ecosystem to deploy autonomy at scale, from robotaxis and... ...investors.About the RoleAs a Principal Software Engineer, you will help define and... ...Performance, and Onboard Systems, requiring deep technical... ...direction of Nuro’s onboard infrastructure. We are looking for a...Immediate startFlexible hours- ...leader in AI cloud infrastructure serving tens of... ...and unite these systems, making groundbreaking... ...Infrastructure Engineering organization... ...seasoned Staff Storage Software Engineer with... ...protocol solutions at scale across object,... ...environments, rack-scale infrastructure...Work at officeLocal areaWork from homeFlexible hours
$272k - $431.25k
...MODS organization seeks a Principal Engineer to architect and scale next-generation L10 and L11 diagnostic systems for Cloud Service Providers... ...systems and hardware / software interfaces is essential for... ...their unique data center infrastructures.What we need to see:Bachelor...Full time- ...framework and library. He/she will participate in the core system design and development. Our target system is based on... ...components. Qualifications ~5+ years proven records on infrastructure level software development experience ~2+ years Clojure development...Full time
$184k - $287.5k
...cutting‑edge hardware and software innovation to deliver... ...group of forward‑thinking engineers tackling some of the globe... ...re searching for a Senior Systems Software Engineer with deep... ...problems at large scale and help shape how AI infrastructure runs in production. In this...Full timeRemote work$248k - $406k
...business needs of the team.The Systems and Infrastructure Team is building a next-generation engineering system and infrastructure... ...to re-platform the largest GPU-scale inference and training infrastructure... ...one of the most leveraged software problems at the company and build...For contractorsWork experience placementWork at officeFlexible hours$253k - $416k
...transform the way the world works.Job DescriptionDistinguished Software Engineer, Systems Infrastructure - Compute, Deployment, Infrastructure as Code (At... ..., which provides shared system services within a large-scale, private, on-premise distributed environment.This role...For contractorsWork experience placementWork at officeFlexible hours$144k - $236k
...DescriptionAs part of our world-class software engineering team, you will help build the next-generation infrastructure and platforms that power... ...data infrastructure, storage systems, streaming platforms, traffic... ...operate AI-first products at scale. You will work on...For contractorsWork experience placementWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!
- principal software engineer Santa Clara, CA
- systems engineer Santa Clara, CA
- system performance engineer Santa Clara, CA
- software system engineer Santa Clara, CA
- computer system validation engineer Santa Clara, CA
- healthcare systems engineer Santa Clara, CA
- distributed systems engineer Santa Clara, CA
- advanced systems engineer Santa Clara, CA
- sr systems engineer Santa Clara, CA
- senior windows systems engineer Santa Clara, CA


