Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure (Santa Clara)

$272k - $431.25k
Full-time

NVIDIA Corporation

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.

What You Will Be Doing:

  • Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user‑mode manageability software.
  • Use Kubernetes and cloud‑native primitives as an infrastructure fabric when appropriate, including controllers, operators, reconciliation loops, and open source components.
  • Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.
  • Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces.
  • Translate forward‑looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.
  • Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real‑world deployment and integration needs.
  • Establish reliability, security, validation, and left‑shift strategies that reduce risk before hardware reaches production environments.
  • Mentor senior engineers and technical leads, raising the engineering bar for large‑scale networked systems, foundational software, and rack‑scale control plane development.
  • Make high‑quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long‑term infrastructure evolution.

What We Need To See:

  • BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.
  • Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.
  • Practical coding skills in Go, C++, or Rust, including the capability to write, review, and direct production‑quality infrastructure software. Experience with Rust is highly valued.
  • Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large‑scale infrastructure services.
  • Experience with Linux‑based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring‑up workflows.
  • Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric‑level manageability.
  • Experience with complex accelerator‑based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high‑performance computing systems.
  • Expertise in in‑band and out‑of‑band management architectures, including BMCs, Redfish, IPMI, and related system management protocols.
  • Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.
  • Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment‑specific integrations.
  • Experience using AI‑assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.
  • Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations.
  • Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.

Ways To Stand Out from the Crowd:

  • Built software supporting multiple adoption models — internal services, CSP‑integrated offerings, reusable libraries, and customer‑extensible APIs.
  • Strong Rust skills in systems, infrastructure, or hardware‑adjacent software.
  • Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI‑assisted engineering.
  • Hands‑on with fleet‑scale provisioning, updates, rollback, observability, health, and remediation.
  • Led across the full data center product lifecycle: inception, pre‑ and post‑silicon, manufacturing, deployment, and operations.
  • Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.
  • Deep experience with rack‑ or cluster‑scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain.
  • Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long‑term direction.

Compensation and Benefits:

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000USD– 431,250USD. You will also be eligible for equity and benefits.

Equal Employment Opportunity:

NVIDIA is committed to fostering a diverse work environment and is proud to be an equal opportunity employer. We highly value diversity in our current and future employees, and we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

#J-18808-Ljbffr
Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure (Santa Clara) in Santa Clara, CA vacancy
  • $272k - $431.25k

     ...team and see how you can make a lasting impact on the world. At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and... 
    Suggested
    Full time
    Shift work

    NVIDIA

    Santa Clara, CA
    4 hours ago
  • $272k - $425.5k

    Principal Software Engineer – Large-Scale LLM Memory and Storage Systems page is loaded## Principal Software Engineer – Large-Scale...  ...Systemslocations: US, CA, Santa Clara: US, WA, Remote: US, MA, Remotetime...  ...storage, or ML systems infrastructure in C/C++ and Python, with a... 
    Suggested
    Full time
    Local area
    Remote work

    NVIDIA Corporation

    Santa Clara, CA
    4 hours ago
  •  ...NVIDIA Corporation is seeking a Principal Rack Scale Systems Infrastructure Engineer in Santa Clara, California. In this role, you will define software architecture for rack-scale infrastructure products and mentor engineers while collaborating with hardware teams. The... 
    Suggested
    Full time

    NVIDIA Corporation

    Santa Clara, CA
    4 hours ago
  •  ...NVIDIA is seeking a Sr. Principal Systems Software Engineer in Santa Clara to lead development of GPU-accelerated data processing for the Apache...  ...optimize performance and scalability at large scale, shaping infrastructure, CI, and testing strategies while #J-18808-Ljbffr
    Suggested
    Full time

    NVIDIA

    Santa Clara, CA
    4 hours ago
  •  ...Secure Cloud and AI infrastructure is the...  ...seeking a world‑class Principal Engineer (Sr Manager‑equivalent...  ...standards for software quality, and...  ...efficiency at a global scale, we want to hear...  ...at our dynamic Santa Clara California...  ...implementation of novel systems that leverage... 
    Suggested
    Full time
    Work at office
    3 days per week

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

    Principal System Software Engineer - Data Center MODS page is loaded## Principal System...  ...MODSlocations: US, CA, Santa Clara: US, CA, Remotetime type:...  ...Engineer to architect and scale next-generation L10 and L...  ...their unique data center infrastructures.**What we need to see:***... 
    Full time

    NVIDIA Corporation

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

     ...Overview NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration...  ...processing problems challenges at large scale Provide recommendations and...  ...decisions surrounding topics such as infrastructure, continuous integration and testing... 
    Full time
    Work experience placement

    NVIDIA

    Santa Clara, CA
    4 hours ago
  • $208k - $260k

     ...manage their hybrid cloud infrastructure. Gigamon has served...  ...organizations. As a Principal Software Engineer on the Network Management System team, you will lead the...  ...is based out of our Santa Clara, CA headquarters, following...  ...that support large-scale deployments and long-... 
    Local area
    Worldwide
    3 days per week

    Gigamon

    Santa Clara, CA
    22 days ago
  •  ...NVIDIA seeks a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group to drive GPU-accelerated data processing in production...  ...collaborating with OSS communities and distributed systems teams at scale. A strong background in distributed systems and open... 
    Full time

    Nvidia Corporation

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team builds the foundational software that turns...  ..., monitoring, and managing workloads at production scale. Drive triage of the most difficult sequencing, initialization... 
    Full time
    Shift work

    NVIDIA Gruppe

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Principal Software Engineer for LPX System Software to build foundational software for a novel computing architecture in Santa Clara, California. The role involves shaping key system software components in Rust and driving architecture design while... 
    Full time

    NVIDIA

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

     ...productivity required for strong scaling for HPC and generative AI...  ...are looking for expert engineers to come and help design rack level solutions for next...  ...solutions for scaling AI infrastructure using GPUs and Grace...  ...space complexity and project system resource requirements.... 
    Full time

    NVIDIA

    Santa Clara, CA
    4 hours ago
  •  ...Apply for Lead Principal Platform Software Engineer at Oracle in Santa Clara, CA, US. This Full time on site...  ...growth. Oracle Cloud Infrastructure (OCI) Networking is...  ...workloads at global cloud scale. We’re rearchitecting...  ...‑scale distributed systems to join our technical... 
    Full time

    Oracle

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

     ...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations in automotive platform software, system...  ...to the architecture, development, optimization, and scaling of foundational software technologies powering NVIDIA automotive... 
    Full time

    NVIDIA

    Santa Clara, CA
    4 hours ago
  • $224k - $356.5k

     .... We are the Datacenter System Software team, and we are looking for...  ...highly motivated, creative Engineering Manager to drive Factory...  ...This includes tightly coupled rack‑scale systems such as GB200/GB300...  .... Build the end‑to‑end infrastructure and workflows that ensure every... 
    Full time
    Shift work

    NVIDIA

    Santa Clara, CA
    4 hours ago
  •  ...Oracle in Santa Clara is seeking engineers to apply formal specification and verification to cloud-scale distributed systems. The role focuses on practical application of formal methods,...  ...CS and 5+ years of concurrent/distributed software experience. #J-18808-Ljbffr
    Full time

    Oracle

    Santa Clara, CA
    4 hours ago
  •  ...the Layer-7 Security Software team, we are responsible...  ...Content Inspection Engine runs on hardware, virtualized...  ...Layer‑7 security infrastructure. We design and develop...  ...next‑generation firewall system Deliver features...  ...programming and large‑scale, distributed, and/or high... 
    Full time

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

    We are hiring senior engineers to work on the CUDA driver, a core component of our platform...  ...programming model across a range of system configurations and hardware capabilities...  ...experience) 15+ years of relevant systems software development experience Strong C programming... 
    Full time

    NVIDIA

    Santa Clara, CA
    4 hours ago
  • $147k - $237.5k

     ...Layer 7 security team is seeking a Senior Principal Software Engineer to lead the design and development of...  ...This strategic role drives innovation in systems that operate at massive scale and protect mission-critical infrastructure worldwide. You will define architecture,... 
    Full time
    Worldwide

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 hours ago
  • $184k - $287.5k

     ...Senior Software Engineer, Cloud-Native Stack – CSP...  ...locations US, CA, Santa Clara US, TX, Austin US...  ...developing advanced multi-rack, multi-tenant AI/...  ...debugging large-scale, cloud-native...  ...), and infrastructure-as-code. ~ Excellent...  ...experience in distributed systems (Go, Rust, C/C++... 
    Full time

    NVIDIA Corporation

    Santa Clara, CA
    4 hours ago
  • $208k - $260k

     ...their hybrid cloud infrastructure. Gigamon has...  ...organizations.    As a Principal Security Detections Engineer on the GigaSMART...  ...high-performance systems software to identify...  ...based out of our Santa Clara, CA headquarters,...  ...with strong focus on scale, accuracy, and performance... 
    Local area
    Worldwide
    3 days per week

    Gigamon

    Santa Clara, CA
    22 days ago
  •  ...outcomes. Job Summary As a Principal Software Engineer within the Engineering team, you...  ...implement, and troubleshoot high‑scale distributed systems, playing a pivotal role in shaping...  ...microservice architectures, global network infrastructure, and load balancing. Working... 
    Full time
    Work at office

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 hours ago
  • $192k - $304.75k

     ...Scientist with a focus in System Software and I/O! NVIDIA is seeking...  ...programming, programming large‑scale clusters, and experience in...  ...research, hardware engineering, and product groups. Publish...  ...architecture research, software infrastructure development and evaluation.... 
    Full time
    Work experience placement

    NVIDIA

    Santa Clara, CA
    4 hours ago
  •  ...NVIDIA Gruppe in Santa Clara is seeking a Senior Architect to lead projects in AI infrastructure. You will architect and implement high-performance communication libraries...  ...-edge technologies in networking and systems software. The ideal candidate has over 12 years of... 
    Full time

    NVIDIA Gruppe

    Santa Clara, CA
    4 hours ago
  • $272k - $431.25k

     ...Rubin–class compute platform engineered for low‑Earth orbit mission...  ...to own end‑to‑end system software architecture for Space‑1 and...  ...experience in building AI infrastructure and systems in space. Proven...  ...platform software for large‑scale data centers or mission‑critical... 
    Full time
    Work experience placement
    Remote work

    NVIDIA

    Santa Clara, CA
    4 hours ago
  •  ...Senior Solutions Architect located in Santa Clara, California, to work with its Cloud Partners...  ...AI supercomputers and enterprise AI infrastructures. The role requires extensive experience...  ...complex issues within GPU and HPC systems. Successful applicants will also be encouraged... 
    Full time

    NVIDIA

    Santa Clara, CA
    4 hours ago
  •  ...within NVIDIA’s Networking Systems & Software Architecture group is solving some of AI’s hardest infrastructure problems. The team builds systems...  ...research and production engineering! What you will be doing:...  ...about building large‑scale, high‑impact data platforms... 
    Full time

    NVIDIA Gruppe

    Santa Clara, CA
    4 hours ago
  • $99.6k - $234.6k

     ...Job Description The Principal AI Agent / ML Software Engineer is a Senior Staff-level, hands‑on technical...  ...operating next‑generation AI systems on Oracle Cloud Infrastructure (OCI). This person will set...  ...AI applications used in large‑scale, business‑critical... 
    Full time
    Temporary work
    Flexible hours

    Oracle

    Santa Clara, CA
    4 hours ago
  • $184k - $356.5k

     ...Senior System Software Engineer Platform - OpenBMC page is loaded Senior System Software Engineer Platform...  ...- OpenBMC Apply locations US, CA, Santa Clara US, Remote time type Full time posted...  ...platform security for x86/ARM based Rack/Blade server systems. NVIDIA is... 
    Full time
    Second job
    Remote work

    NVIDIA Corporation

    Santa Clara, CA
    4 hours ago
  • $149.41k - $190.46k

     ...Senior Energy Systems Analyst Join to apply for the Senior...  ...electric service to the City of Santa Clara residents and businesses...  ...system expansion and aging infrastructure replacement plan....  ...work in Computer Science, Engineering, Mathematics, or a related... 
    Permanent employment
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Weekend work

    A Hiring Company

    Santa Clara, CA
    4 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure (Santa Clara). Be the first to apply!