Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure

$272k - $431.25k

NVIDIA

Principal Rack Scale Systems Infrastructure Engineer

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology—and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.

What You Will Be Doing:

  • Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.
  • Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.
  • Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.
  • Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.
  • Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.
  • Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.

What We Need To See:

  • BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.
  • Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.
  • Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.
  • Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.
  • Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.
  • Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.
  • Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.
  • Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.
  • Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.

Ways To Stand Out from the crowd:

  • Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.
  • Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.
  • Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.
  • Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 19, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Vacancy posted 20 hours ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure in Santa Clara, CA vacancy
  • $272k - $431.25k

     ...accelerators feel like a single system at datacenter scale. As large language models rapidly...  ...LLM workloads. We are seeking a Principal Systems Engineer to define the vision and roadmap...  ...performance storage, or ML systems infrastructure in C/C++ and Python, with a track... 
    Suggested
    Local area
    Remote work

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $226k - $369k

     ...part of our world-class software engineering team, you will take...  ...the next-generation infrastructure and platforms for LinkedIn...  ..., API design and systems design, and your...  ...performs at massive scale. LinkedIn has pioneered...  ...our company. As a Principal Staff Software Engineer... 
    Suggested
    For contractors
    Work at office
    Flexible hours

    LinkedIn

    Sunnyvale, CA
    4 days ago
  • $248k - $391k

     ...seeking a highly skilled Principal Software Engineer to jo in our dynamic...  ...the performance of our infrastructure both on-prem and in the...  ...internal AI inference platform scaling to frontier-class models...  ...for pre-release, rack-scale GPU systems (including Blackwell and... 
    Suggested

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $230k - $360k

     ...Lead Infrastructure And Reliability Engineer (Systems & Scale) Our Infrastructure Engineering team is a systems engineering group with company-level responsibility. At Luma, reliability engineers work directly with the researchers and products pushing the limits of... 
    Suggested

    Luma AI

    Palo Alto, CA
    4 days ago
  • $207k - $300k

    Google Inc. is looking for a Staff Software Engineer for AI and Infrastructure to contribute to Google Cloud's mission. The ideal candidate will have deep...  ...include designing and implementing computer systems, collaborating on impactful projects, and providing technical... 
    Suggested

    Google Inc.

    Sunnyvale, CA
    2 days ago
  •  ...Principal Software Engineer, Systems/Solutions Test This role has been designed as 'Hybrid' with an expectation...  ...frameworks for functional, scale, resiliency, and performance testing...  .... Strong understanding of cloud infrastructure, core systems, and WAN architectures... 
    Work at office
    2 days per week

    Hewlett Packard Enterprise

    Sunnyvale, CA
    1 day ago
  • $224k - $356.5k

     ...NVIDIA is hiring engineers to scale up the introduction of next...  ...architecture into its EDA Infrastructure. We expect you to have...  ...introductions (NPIs), distributed systems, familiarity with software testing and deployment,...  ...to join the EDA Team. Principal Software Engineer... 

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...Secure Cloud and AI infrastructure is the foundation of...  ...seeking a world-class Principal Engineer (Sr Manager-equivalent...  ...elevate our standards for software quality, and unlock...  ...at a global scale, we want to hear from...  ...implementation of novel systems that leverage Large Language... 
    Full time
    Work at office
    3 days per week

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  •  ...Principal AI/ML System Software Engineer At d-Matrix, we are focused on unleashing the potential of generative...  ...design. You are able to build and scale software deliverables in a tight...  ...to build out the deployment infrastructure, working closely with other software... 
    Work experience placement
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $141k - $202k

    A leading technology company in Sunnyvale is seeking a Software Engineer III to develop infrastructure solutions. The role involves programming in C++, tackling large-scale systems challenges, and ensuring software quality through testing and debugging. Candidates should... 
    Full time

    Google Inc.

    Sunnyvale, CA
    3 days ago
  • CrowdStrike, Inc. is seeking a Senior Infrastructure Engineer based in Sunnyvale, California, to help expand the Falcon platform...  ...role entails designing and implementing large-scale Kubernetes solutions and optimizing system reliability. The ideal candidate will possess... 
    Remote job

    Koitecc Solutions

    Sunnyvale, CA
    3 days ago
  •  ...leading technology company is seeking a Software Engineer to develop next-generation...  ...debugging complex issues across large-scale systems. Candidates should have a strong foundation...  ...dynamic team at the forefront of AI and infrastructure innovation. #J-18808-Ljbffr Google... 

    Google Inc.

    Sunnyvale, CA
    20 hours ago
  •  ..., Security Automation Engineer Walmart is seeking...  ...our endpoint security infrastructure. This role will focus...  ...Infrastructure as Code, automate system reporting, and work...  ...administration, and software development...  ...security at enterprise scale. Our team supports critical... 

    Walmart

    Sunnyvale, CA
    4 days ago
  • $272k - $431.25k

     ...MODS organization seeks a Principal Engineer to architect and scale next-generation L10 and L11 diagnostic systems for Cloud Service Providers...  ...systems and hardware / software interfaces is essential for...  ...their unique data center infrastructures. What we need to see:... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $147k - $211k

    Software Engineer III, Embedded Systems/Firmware, Platforms Infrastructure Engineering Apply X Note: By applying to this position you will have an opportunity to share...  ...delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our... 
    Full time
    Worldwide

    Google Inc.

    Sunnyvale, CA
    4 days ago
  • $184k - $287.5k

     ...the world. Join NVIDIA's software infrastructure team to design, build, and improve software systems for rack, networking, and datacenter...  ...management. As a Senior Software Engineer - Datacenter Systems, you...  ...technology supporting large-scale GPU clusters connected... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...impact on the world. We are seeking a Senior Systems Software Engineer to join our advanced infrastructure software team. In this role, you will be responsible...  ..., developing, and maintaining high-performance, rack-scale management solutions for datacenter environments.... 

    NVIDIA

    Santa Clara, CA
    20 hours ago
  •  ...data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation...  ...ROLE: AMD is looking for a strategic software engineering lead who is passionate about improving...  ...Develop techniques for optimizing scale-up and scale-out inference. Develop... 

    Advanced Micro Devices , Inc.

    Santa Clara, CA
    2 days ago
  • A leading tech company is seeking a Staff Software Engineer in Sunnyvale. The role requires extensive experience in software development and managing large-scale infrastructure. Key responsibilities include providing technical leadership, mentoring engineers, and delivering... 

    Google Inc.

    Sunnyvale, CA
    4 days ago
  • $138.8k - $190.85k

     ...ago Requisition ID: 1659 Sr. Engineer, Infrastructure Software Job Summary SiTime is...  ...Michigan. You will build and scale next‑generation mixed‑signal semiconductor test systems and automation platforms, collaborating...  ...Test Equipment (ATE), rack‑level test infrastructure... 
    Full time
    Overseas

    SiTime Corporation

    Santa Clara, CA
    20 hours ago
  • System / Clojure Principal Software Engineer Integrated Resources, Inc is a premier staffing firm recognized as one of the tri-state's most well-respected...  ...components. Responsibilities: Develop and maintain infrastructure-level software solutions. Contribute to the design... 

    Integrated Resources Inc.

    Santa Clara, CA
    3 days ago
  • $160.36k - $240.54k

     ...Senior Software Engineer – GenAI Infrastructure & Agent Systems for Engineering Efficiency Mountain View, California (HQ) Who We Are Nuro is a self-driving...  ...platforms a clear path to AVs at commercial scale, empowering a safer, richer, and more connected future... 

    Nuro

    Mountain View, CA
    1 day ago
  • $108k - $162k

     ...in process control, combining global scale with an expanded portfolio of leading...  ...We are seeking a highly skilled Sr. Systems & Infrastructure Engineer to join a dynamic, security-first IT...  ...buildouts, hardware refresh planning, rack/power design, and operational support... 
    Permanent employment

    Onto

    Milpitas, CA
    3 days ago
  • $165k - $242k

     ...Senior Software Engineer, Data Center Infrastructure Tooling CoreWeave is The Essential Cloud...  ...innovators to build and scale AI with confidence. Trusted...  ...relationships across racks, rows, and floors. The schema...  ...with internal/external systems and data sources that feed... 
    Temporary work
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • $248k - $406k

     ...We are seeking a Distinguished Engineer to help shape the future of LinkedIn's Core Infrastructure platform, which provides shared system services within a large-scale, private, on-premise distributed...  ...for quality and efficiency of software systems while balancing business... 
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    LinkedIn

    Mountain View, CA
    2 days ago
  • $272k - $431.25k

     ...productivity required for strong scaling for HPC and generative AI...  ...are looking for expert engineers to come and help design rack level solutions for next...  ...solutions for scaling AI infrastructure using GPUs and Grace...  ...space complexity and project system resource requirements.... 

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $349k - $431k

     ...Principal Software Engineer, ML System Architect Waymo is an autonomous driving technology company with the...  ...are increasingly leveraging large-scale Foundation Models to unlock new capabilities...  ..., deep learning frameworks, and AI infrastructure. ~ A track record of architecting... 
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $249k

     ...more open world. Join us. Principal Software Development Engineer — Platform & Infrastructure Introduction to Team:...  ...leader who can design complex systems, implement solutions in production...  ...traffic engineering for high-scale workloads Own and evangelize... 
    Local area
    Flexible hours

    Expedia Group

    San Jose, CA
    20 hours ago
  • $141k - $202k

    A leading technology company in California is seeking a Software Engineer to develop innovative technologies that change how users connect...  ...interact. The ideal candidate will collaborate on large-scale infrastructure and AI solutions, requiring programming experience in C++... 

    Google Inc.

    Sunnyvale, CA
    1 day ago
  •  ...leading technology company is seeking software engineers for its Sunnyvale, CA office. Responsibilities...  ...Controller and optimizing Google's infrastructure. Candidates should have a Bachelor’s...  .... Experience with distributed systems is preferred. The role offers a competitive... 
    Work at office

    Google Inc.

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!