Principal Software Engineer - Rack Scale Systems Infrastructure

$272k - $431.25k

Full-time

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally. What You Will Be Doing: Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software. Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments. Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization. Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments. Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development. Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution. What We Need To See: BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering. Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs. Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued. Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows. Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems. Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation. Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations. Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation. Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives. Ways To Stand Out from the crowd: Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software. Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation. Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs. Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until July 2, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

Apply

Vacancy posted 6 hours ago

Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure in Santa Clara, CA vacancy

Principal Software Engineer - Rack Scale Systems Infrastructure
$272k - $431.25k
...team and see how you can make a lasting impact on the world. At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and...
Suggested
Shift work
Jobleads-US
Santa Clara, CA
3 days ago
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure... ...computing. Today, our AI infrastructure powers global intelligence, transforming...
Suggested
Full time
Shift work
NVIDIA
Santa Clara, CA
2 days ago
Principal Rack-Scale System Software Engineer
$272k - $431.25k
NVIDIA Corporation is seeking a Principal Software Engineer to join the CSP Engagements team in Santa Clara, CA. This role is pivotal for driving rack-scale system software and firmware architecture, ensuring seamless integration, operation, and monitoring of systems at...
Suggested
NVIDIA Corporation
Santa Clara, CA
1 day ago
Principal Software Engineer - Large-Scale LLM Memory and Storage Systems
$272k - $425.5k
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems page is loaded## Principal Software Engineer – Large-Scale LLM Memory and Storage Systemslocations... ..., high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track record of...
Suggested
Local area
Remote work
NVIDIA Corporation
Santa Clara, CA
3 days ago
Senior Systems Engineer — Rust/Go/C++ for Rack-Scale Infra
NVIDIA Corporation is seeking a Senior Systems Software Engineer to join its advanced infrastructure software team in Santa Clara, California. You will be responsible... ..., developing, and maintaining high-performance, rack-scale management solutions. The role emphasizes work in...
Suggested
NVIDIA Corporation
Santa Clara, CA
5 days ago
Senior Infrastructure Software Engineer - Large-Scale Systems
Google Inc. in Sunnyvale, CA is looking for a Software Engineer to develop next-generation technologies crucial to Google... ...ideal candidate will have experience with large-scale infrastructure and distributed systems, along with proficiency in programming languages such...
Google Inc.
Sunnyvale, CA
3 days ago
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements... ...point for fleet-scale reliability, working... ...you to distinguish systemic architectural gaps... ...expertise in multi-NUMA, rack-scale system... ...accelerator or HPC infrastructure Background in defining...
NVIDIA Gruppe
Santa Clara, CA
1 day ago
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements... ...firmware and GPU system software, working... ...firmware at fleet scale. You will drive work... ...hundreds of GPUs per rack Serve as the technical... .... Today, our AI infrastructure powers global intelligence...
Full time
NVIDIA
Santa Clara, CA
2 days ago
Staff AI & Infrastructure Engineer — Scale & Systems
$207k - $300k
Google Inc. is looking for a Staff Software Engineer for AI and Infrastructure to contribute to Google Cloud's mission. The ideal candidate will have deep... ...include designing and implementing computer systems, collaborating on impactful projects, and providing technical...
Google Inc.
Sunnyvale, CA
2 days ago
Principal Software Engineer, Systems/Solutions Test
$172k - $349k
## Principal Software Engineer, Systems/Solutions TestApplylocations: Sunnyvale, California, United States... ...automation frameworks for functional, scale, resiliency, and performance... ...patterns.* Strong understanding of cloud infrastructure, core systems, and WAN...
Work experience placement
Work at office
2 days per week
Jobleads-US
Sunnyvale, CA
4 days ago
Software Engineer III, Embedded Systems/Firmware, AI and Infrastructure
$147k - $211k
Software Engineer III, Embedded Systems/Firmware, AI and Infrastructure Google, Sunnyvale, CA, USA Qualifications Bachelor’s degree or equivalent practical experience... ...2 years of experience with performance, large scale systems data analysis, visualization tools, or debugging...
Google Inc.
Sunnyvale, CA
5 days ago
Senior Software Engineer, Embedded Systems/Firmware, AI and Infrastructure
$174k - $252k
Senior Software Engineer, Embedded Systems/Firmware, AI and Infrastructure Sunnyvale, CA, USA Bachelor’s degree or equivalent practical experience. 5 years of experience... ...products need to handle information at massive scale, and extend well beyond web search. We're looking...
Full time
Worldwide
Google Inc.
Sunnyvale, CA
2 days ago
Senior AI Infrastructure Engineer, Large-Scale GPU Clusters
...Corporation in Santa Clara is seeking a Senior Software Engineer to lead the optimization of large-scale AI systems. This role will involve profiling and tuning... ...will have over 8 years of experience in software infrastructure for AI systems, with expert-level programming...
NVIDIA Corporation
Santa Clara, CA
3 days ago
Senior Infrastructure Engineer, Large-Scale AI & Platform
...leading technology company is seeking a Software Engineer to develop next-generation... ...debugging complex issues across large-scale systems. Candidates should have a strong foundation... ...dynamic team at the forefront of AI and infrastructure innovation. #J-18808-Ljbffr Google...
Google Inc.
Sunnyvale, CA
5 days ago
Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud
...cutting‑edge hardware and software innovation to deliver... ...a team of innovative engineers dedicated to solving... ...an outstanding Senior Systems Software Engineer with... ...real‑world problems at scale. What You’ll Be Doing:... ...one of public CSP infrastructure (GCP, AWS, Azure, OCI,...
Worldwide
NVIDIA Corporation
Santa Clara, CA
3 days ago
Staff Software Engineer, Cloud Infrastructure — Scale & AI
$207k - $301k
Google Inc. is seeking a Staff Software Engineer for its Infrastructure team in Sunnyvale, CA. This pivotal role focuses on developing advanced technologies within large-scale systems, distributed computing, and AI infrastructure, enabling billions to connect and interact...
Google Inc.
Sunnyvale, CA
3 days ago
Senior Principal Software Engineer ( Cloud Infrastructure and Platform Engineering )
...Secure Cloud and AI infrastructure is the foundation of... ...seeking a world‑class Principal Engineer (Sr Manager‑equivalent... ...elevate our standards for software quality, and unlock... ...at a global scale, we want to hear from... ...implementation of novel systems that leverage Large Language...
Full time
Work at office
3 days per week
Palo Alto Networks
Santa Clara, CA
2 days ago
L11 Rack-Scale AI System Product Engineer (DGX/MGX)
$168k - $258.75k
NVIDIA AI in Santa Clara is seeking a highly motivated System Product Development Engineer to lead the development and productization of DGX and MGX products. You will drive the integration of advanced AI supercomputing systems, focusing on quality and efficiency. The role...
NVIDIA AI
Santa Clara, CA
5 days ago
L11 Rack-Scale AI System Development Engineer
$168k - $258.75k
NVIDIA Corporation is seeking a System Product Development Engineer in Santa Clara, California, to lead the development and productization of DGX and MGX products. The role involves driving integration efforts, collaborating with various teams, and ensuring optimized designs...
NVIDIA Corporation
Santa Clara, CA
1 day ago
Senior Systems Software Engineer - Rust, Go, C++
...Senior Systems Software Engineer – Advanced Infrastructure Software Team We are seeking a Senior Systems Software Engineer to join our advanced infrastructure... ..., developing, and maintaining high-performance, rack-scale management solutions for datacenter environments....
NVIDIA
Santa Clara, CA
1 day ago
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
$272k - $431.25k
...We're looking for a Principal Engineer to join our CSP Engagements... ...patterns and drive systemic improvements in... ...the latest NVIDIA rack-scale systems, GPU architectures... ...configuration, software, or workload differences... ...work (profiling infrastructure, benchmark harnesses...
NVIDIA Corporation
Santa Clara, CA
1 day ago
Staff Software Engineer, Embedded Systems/Firmware, Platforms Infrastructure Engineering
$207k - $301k
Staff Software Engineer, Embedded Systems/Firmware, Platforms Infrastructure Engineering Google, Sunnyvale, CA, USA Requirements Bachelor's degree or equivalent practical... ...products need to handle information at massive scale, and extend well beyond web search. We look for...
Google Inc.
Sunnyvale, CA
3 days ago
Principal System Software Engineer - Data Center MODS
$272k - $431.25k
Principal System Software Engineer - Data Center MODS page is loaded## Principal System Software Engineer... ...Principal Engineer to architect and scale next-generation L10 and L11... ...challenges within their unique data center infrastructures.**What we need to see:*** Bachelor'...
Jobleads-US
Santa Clara, CA
2 days ago
Senior Systems Software Engineer - GPU Performance at Scale
$184k - $287.5k
...world. We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely... ...performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies...
Remote work
NVIDIA
Santa Clara, CA
2 days ago
Sr. Engineer, Infrastructure Software
$138.8k - $190.85k
...ago Requisition ID: 1659 Sr. Engineer, Infrastructure Software Job Summary SiTime is... ...Michigan. You will build and scale next‑generation mixed‑signal semiconductor test systems and automation platforms, collaborating... ...Test Equipment (ATE), rack‑level test infrastructure...
Full time
Overseas
SiTime Corporation
Santa Clara, CA
5 days ago
Software Engineering Manager II, Embedded Systems/Firmware, Platforms Infrastructure Engineering
$207k - $301k
Software Engineering Manager II, Embedded Systems/Firmware, Platforms Infrastructure Engineering Google Sunnyvale, CA, USA Requirements Bachelor's degree or equivalent practical... ...processing, distributed computing, large‑scale system design, networking, security, data...
Google Inc.
Sunnyvale, CA
2 days ago
System / Clojure Principal Software Engineer
System / Clojure Principal Software Engineer Integrated Resources, Inc is a premier staffing firm recognized as one of the tri-state's most well-respected... ...components. Responsibilities: Develop and maintain infrastructure-level software solutions. Contribute to the design...
Integrated Resources
Santa Clara, CA
3 days ago
Senior Software Engineer, Data Center Infrastructure Tooling
$165k - $242k
...Senior Software Engineer, Data Center Infrastructure Tooling CoreWeave is The Essential Cloud... ...innovators to build and scale AI with confidence. Trusted... ...relationships across racks, rows, and floors. The schema... ...with internal/external systems and data sources that feed...
CoreWeave
Sunnyvale, CA
4 days ago
IT Infrastructure Systems Engineer
$108k - $162k
...in process control, combining global scale with an expanded portfolio of leading... ...We are seeking a highly skilled Sr. Systems & Infrastructure Engineer to join a dynamic, security-first IT... ...buildouts, hardware refresh planning, rack/power design, and operational support...
Permanent employment
Onto
Milpitas, CA
3 days ago
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing... ...of performance practices in large‑scale GPU infrastructure, delivering powerful tools, methodologies, and flows...
NVIDIA Corporation
Santa Clara, CA
4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!