Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure

$272k - $431.25k

NVIDIA

US, CA, Santa Clara

US, NC, Remote

US, TX, Remote

US, Remote

US, MA, Remote

Full time

JR2017966

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.

What You Will Be Doing:

  • Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.

  • Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.

  • Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.

  • Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.

  • Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.

  • Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.

What We Need To See:

  • BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.

  • Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.

  • Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.

  • Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.

  • Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.

  • Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.

  • Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.

  • Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.

  • Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.

Ways To Stand Out from the crowd:

  • Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.

  • Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.

  • Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.

  • Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits ( .

Applications for this job will be accepted at least until September 14, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry.

Learn more about NVIDIA .

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure in Santa Clara, CA vacancy
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role... 
    Suggested
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    7 days ago
  • $272k - $431.25k

     ...Principal Systems Engineer NVIDIA Dynamo is a high-throughput, low-latency inference framework for...  ...feel like a single system at datacenter scale. As large language models rapidly...  ...high-performance storage, or ML systems infrastructure in C/C++ and Python, with a track... 
    Suggested
    Local area
    Remote work

    NVIDIA

    Santa Clara, CA
    4 days ago
  •  ...are at the forefront of software and hardware innovation,...  ...per week. The role: Principal System Software Engineer, AI Inference Execution...  ...You are able to build and scale software deliverables in...  ...build out the deployment infrastructure, working closely with other... 
    Suggested
    Work experience placement
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    2 days ago
  • $114.6k - $234.6k

     ...Principal Systems Software Engineer Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems...  ...of software, firmware, hardware, and large-scale cloud infrastructure. You will work on complex GPU... 
    Suggested
    Temporary work
    Flexible hours

    Hackajob

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...JR2021260 NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration...  ...processing problems challenges at large scale Provide recommendations and...  ...decisions surrounding topics such as infrastructure, continuous integration and testing... 
    Suggested
    Full time
    Work experience placement

    NVIDIA

    Santa Clara, CA
    3 days ago
  •  ...Google in Sunnyvale, CA is hiring a Software Engineering Manager to lead multiple teams working on large-scale infrastructure and distributed systems. You will guide project goals, help set product strategy, and mentor engineers across locations while overseeing budgets... 

    Jobleads-US

    Sunnyvale, CA
    4 days ago
  • $248k - $391k

     ...seeking a highly skilled Principal Software Engineer to join our dynamic...  ...performance of our infrastructure both on-prem and in...  ...-class AI inference systems. Join us in this...  ...inference platform scaling to frontier-class models...  ...for pre-release, rack-scale GPU systems (including... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation...  ...architecture, development, optimization, and scaling of foundational software...  ...accelerated computing. Today, our AI infrastructure powers global intelligence, transforming... 
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...Secure Cloud and AI infrastructure is the foundation of...  ...seeking a world‑class Principal Engineer (Sr Manager‑equivalent...  ...elevate our standards for software quality, and unlock...  ...at a global scale, we want to hear from...  ...implementation of novel systems that leverage Large Language... 
    Full time
    Work at office
    3 days per week

    Jobleads-US

    Santa Clara, CA
    1 day ago
  •  ...framework and library. He/she will participate in the core system design and development. Our target system is based on...  ...components. Qualifications ~5+ years proven records on infrastructure level software development experience ~2+ years Clojure development... 
    Full time

    Integrated Resources Inc.

    Santa Clara, CA
    20 hours ago
  •  ...methods.  Job Summary: The AI Server/Rack System Engineer leads the end-to-end integration,...  ...with electrical, mechanical, software, thermal, and R&D teams to drive timely...  ...years of experience with data center infrastructure, enterprise servers, AI server/rack systems... 
    Local area

    Foxconn-PCE Technology

    Santa Clara, CA
    10 days ago
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team builds the foundational software that turns...  ..., monitoring, and managing workloads at production scale. Drive triage of the most difficult sequencing, initialization... 
    Shift work

    NVIDIA Gruppe

    Santa Clara, CA
    20 hours ago
  • $226k - $369k

     ...status and approval. LinkedIn’s Reliability Infrastructure team is responsible for defining and...  ...that keep LinkedIn’s most critical systems stable, resilient, and available at massive scale.As a Principal Staff Software Engineer, Reliability Infrastructure, you will serve... 
    For contractors
    Work at office
    Remote work
    Work from home
    Flexible hours

    Linkedin

    Mountain View, CA
    5 days ago
  • $142.8k - $274.8k

     ...Less than 25%Profession: Software EngineeringDiscipline: Software...  ...: MicrosoftOverviewThe AI Infrastructure Engineering Systems team at Microsoft builds...  ...operate AI services at scale. We create foundational capabilities...  ...inclusive culture.As a Principal Software Engineer - AI... 
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    6 days ago
  • $184k - $287.5k

     ...world. We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely...  ...performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies... 
    Full time
    Remote work

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $182k - $242k

     ...innovators to build and scale AI with confidence....  ...CoreWeave combines superior infrastructure performance with deep...  ...looking for a Senior Engineer to be a driving force...  ...and observability systems that make our AI infrastructure...  ...and comparable across racks, clusters, and... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    10 days ago
  • $2,000 per month

     ...intelligence. We co-design chips, racks, software, and manufacturing to...  ...and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest...  ...includes building and scaling our hybrid high-...  ...deep understanding of systems. It’s not just about writing... 
    Work at office
    Relocation package

    Etched

    San Jose, CA
    18 days ago
  • $258k - $387k

     ...ecosystem to deploy autonomy at scale, from robotaxis and...  ...About the Role As a Principal Software Engineer, you will help define and...  ...Performance, and Onboard Systems, requiring deep technical...  ...direction of Nuro's onboard infrastructure. We are looking for a technical... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    10 days ago
  • $300 per month

     ...vertically integrated AI infrastructure company built...  ...believe in the scale of our ambition...  ...Center Infrastructure Engineering (DCIE) team is...  ..., provision systems, and observability...  ...skilled and motivated Software Engineer to join...  ...within GPU racks and high-density... 
    Temporary work

    Crusoe

    Sunnyvale, CA
    18 days ago
  • $272k - $431.25k

     ...We're looking for a Principal Software Engineer to join our CSP Engagements team...  ...technical focal point for fleet-scale reliability, working...  ...enables you to distinguish systemic architectural gaps from environmental...  ...expertise in multi-NUMA, rack-scale system software and... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as...  ...point for GPU firmware and GPU system software, working directly...  ...NVIDIA GPU firmware at fleet scale. You will drive work streams...  ...across multiple GPUs in a rack-scale systemUnderstanding of... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    7 days ago
  • $248k - $391k

     ...world.We are looking for a Principal Software Engineer to join our Configuration...  ...the future of enterprise infrastructure automation, configuration...  ...build foundational software systems that manage infrastructure...  ...challenges, and drive large-scale transformations from... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $200k - $220k

     ...electric vehicle, grid infrastructure, industrial HVAC and other...  ...is driving the fast-scale growth of the company’s...  ...manufacturing, we are seeking a Principal Machine Control Software Engineer to support the...  ...maintenance of equipment control systems for our semiconductor... 
    Full time
    Temporary work

    Halo Industries, Inc.

    Santa Clara, CA
    20 hours ago
  • $174k - $253k

    Write and test product or system development code....  ...code developed by other engineers and provide feedback...  ...maintaining, or launching software products, and 1 year...  ...at massive scale, and extend well beyond...  ...built by the Technical Infrastructure team to keep it running... 
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • $114.6k - $234.6k

     ...components of scalable, elastic distributed systems. Defines and enforces scalability...  ...and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms...  ....Only Oracle brings together the data, infrastructure, applications, and expertise to power... 
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  • $114.6k - $234.6k

    Oracle Cloud Infrastructure (OCI) delivers mission-critical...  ...offers unmatched hyper-scale, multi-tenant services...  ...are hoping to enhance engineering efficiency by...  ...on building low level systems with high performance...  ...investment and drive the software design and development... 
    Temporary work
    Worldwide
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  • $146.3k - $306.4k

    Defines architecture for large-scale systems software, firmware integration, and fleet automation...  ..., and cost at hyperscale. Champions engineering excellence: coding standards, threat...  ...Oracle brings together the data, infrastructure, applications, and expertise to power... 
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    6 days ago
  • $224k - $356.5k

    NVIDIA is hiring engineers to scale up the introduction of next generation...  ...into its EDA Infrastructure. We expect you to have...  ...introductions (NPIs), distributed systems, familiarity with software testing and deployment,...  ...to join the EDA Team. Principal Software Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...across clouds, on-premises infrastructure, and devices. Our platform...  ...experts in cryptography, systems, and distributed computing...  .... The Role Staff Software Engineer (Rust) - Confidential Computing...  ...AI workloads at scale. Location: Santa Clara... 
    Full time
    H1b

    Fortanix

    Santa Clara, CA
    20 hours ago
  •  ...company is creating the digital infrastructure needed to bring...  ...infrastructure, operating systems, and autonomy. Eighteen of...  ...operating system.  As a Software Engineer on the team, you will develop...  ...designing and shipping large-scale software or systems ~ Demonstrated... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    20 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!