Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure

$272k - $431.25k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.

What You Will Be Doing:
  • Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.
  • Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.
  • Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.
  • Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.
  • Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.
  • Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.
What We Need To See:
  • BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.
  • Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.
  • Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.
  • Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.
  • Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.
  • Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.
  • Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.
  • Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.
  • Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.
Ways To Stand Out from the crowd:
  • Built software supporting multiple adoption models - internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.
  • Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.
  • Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.
  • Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 14, 2026.

This posting is for an existing vacancy.


NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure in United States vacancy
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role... 
    Suggested
    Full time
    Remote work
    Shift work

    Nvidia

    Austin, TX
    7 days ago
  • $272k - $431.25k

     ...accelerators feel like a single system at datacenter scale. As large language models rapidly...  ...LLM workloads. We are seeking a Principal Systems Engineer to define the vision and roadmap...  ...performance storage, or ML systems infrastructure in C/C++ and Python, with a track... 
    Suggested
    Local area
    Remote work

    NVIDIA

    United States
    17 hours ago
  •  ...deploys frontier compute infrastructure fastest will decide whether...  ...spanning hardware and software. Speed and scale are our key differentiators...  ...to the world. The Rack Systems Team Rack Systems owns the...  ...hardware existed. Run a small engineering team with the review rigor... 
    Suggested

    FluidStack

    Austin, TX
    2 days ago
  • $226k - $369k

     ...approval. We are seeking a Principal Staff Software Engineer to join our organization....  ...-performing at global scale.As a Principal Staff Engineer...  ...driving next-generation infrastructure that powers AI-first...  ...large-scale infrastructure systems for end-to-end software lifecycle... 
    Suggested
    For contractors
    Work at office
    Remote work
    Work from home
    Flexible hours

    Linkedin

    Mountain View, CA
    17 hours ago
  • $114.6k - $234.6k

     ...Description Owns complex system modules and fleet...  ...contracts across software, firmware, and hardware...  ...provisioning pipelines at scale. Elevates code quality...  ...and coaching to engineers to drive improvements....  ...brings together the data, infrastructure, applications, and expertise... 
    Suggested
    Temporary work
    Flexible hours
    Shift work

    Oracle

    Phoenix, AZ
    2 days ago
  •  ...Principal Systems Software Engineer Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems...  ...of software, firmware, hardware, and large-scale cloud infrastructure. You will work on complex GPU... 

    Oracle

    Santa Clara, CA
    17 hours ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Sr. Systems Software Engineer for the Apache Spark Acceleration...  ...challenges at large scale Provide recommendations...  ...topics such as infrastructure, continuous integration and...  ...July 19, 2026. Apply here: Principal Systems Software Engineer
    Work experience placement

    University of Illinois

    Champaign, IL
    2 days ago
  •  ...Hardware Platform Engineer Job Description...  ...and deployable at scale. This role ensures...  ...Lead integration of rack-level hardware...  ...and scale Drive system-level design across...  ...compute, cooling, and infrastructure interfaces...  ...deployment Partner with software/data teams to... 
    Temporary work
    Remote work
    Flexible hours

    hackajob

    United States
    4 days ago
  • $160k - $220k

     ...innovate with new products, software features, and new...  ...given as a gift, so our infrastructure must support rapid scalability...  ...small infrastructure engineering team in our NYC and SF...  ...flexible backend systems for a Rails API. These systems scale reliably while supporting... 
    Full time
    Part time
    Flexible hours

    Aura Home, Inc.

    New York, NY
    1 day ago
  • $114.6k - $234.6k

     ...We are looking for smart systems software engineers with BS/MS/PhD in Computer...  ...X11, and Exadata Expansion Rack. Exadata group also keeps looking...  ...distributed algorithms to scale systems Career Level -...  ...brings together the data, infrastructure, applications, and expertise... 
    Temporary work
    Flexible hours

    Oracle

    Oklahoma City, OK
    1 day ago
  • $251k - $352k

    Senior Principal Engineer, Infrastructure Base pay range $251,000.00/yr - $352,000.00/yr...  ...these foundational systems work together seamlessly to...  ...Architect billing systems that scale to support Docker's growth...  ...Technical Expertise 15+ years of software engineering experience... 
    Full time
    Contract work
    Immediate start
    Remote work

    Docker

    Seattle, WA
    5 days ago
  •  ...of next-generation AI infrastructure, unifying Bare-Metal-as...  ...between silicon and software to optimize the I/O path for massive-scale generative AI workloads...  ...in Computer Science or Engineering and a proven track...  ...Co-Design Distributed Systems Bare-Metal-as-a-Service... 
    Temporary work

    FELICIS

    San Francisco, CA
    2 days ago
  • $260k - $340k

     ...the only vertically integrated AI infrastructure company built from the ground up, we...  ...sense of urgency, who believe in the scale of our ambition and thrive on a...  ...Crusoe. About This Role: As the Principal Systems Software Engineer , you will serve as the visionary... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    23 days ago
  • $258k - $387k

     ...ecosystem to deploy autonomy at scale, from robotaxis and...  ...About the Role As a Principal Software Engineer, you will help define and...  ...Performance, and Onboard Systems, requiring deep technical...  ...direction of Nuro's onboard infrastructure. We are looking for a technical... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    15 days ago
  • $108k - $162k

     ...in process control, combining global scale with an expanded portfolio of leading...  ...We are seeking a highly skilled Sr. Systems & Infrastructure Engineer to join a dynamic, security-first IT...  ...buildouts, hardware refresh planning, rack/power design, and operational support... 
    Permanent employment

    Onto

    Milpitas, CA
    3 days ago
  • $248k - $391k

     ...seeking a highly skilled Principal Software Engineer to join our dynamic...  ...performance of our infrastructure both on-prem and in...  ...-class AI inference systems. Join us in this...  ...inference platform scaling to frontier-class models...  ...for pre-release, rack-scale GPU systems (including... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $114.6k - $234.6k

     ...scalable, elastic distributed systems. Defines and enforces...  ...paths for high throughput, hyper scale workloads; and leverages data...  ...Responsibilities Utilize standard software development practices and...  ...embedded Linux application/infrastructure. Experience in one or more... 
    Temporary work
    Remote work
    Flexible hours
    Nashville, TN
    5 days ago
  • $114.6k - $234.6k

     ...degree in Computer Science, Computer Engineering, or a related discipline, or equivalent...  ...looking for 7+ years of professional software engineering experience with clear impact on large-scale distributed systems or cloud infrastructure. We need strong experience... 
    Full time
    Work at office
    Worldwide
    Relocation
    Relocation package
    Flexible hours

    Oracle

    Nashville, TN
    1 day ago
  •  ...methods.  Job Summary: The AI Server/Rack System Engineer leads the end-to-end integration,...  ...with electrical, mechanical, software, thermal, and R&D teams to drive timely...  ...years of experience with data center infrastructure, enterprise servers, AI server/rack systems... 
    Local area

    Foxconn-PCE Technology

    Santa Clara, CA
    29 days ago
  •  ...manages all applications and next steps. Our partner is looking for a Senior Infrastructure Engineer, Core Systems based in United States. This role offers the opportunity to shape and scale next-generation infrastructure across private and public cloud environments.... 
    Full time
    Remote work
    Flexible hours

    jobgether

    United States
    2 days ago
  • $145k - $165k

     ...Job Description Job Description Title: Senior. Systems Engineer - Infrastructure & Cloud Location: Houston, TX Salary: $145,000 - $165,...  ...Experience designing, implementing, and supporting enterprise-scale Linux server platforms. ~ Extensive experience... 
    Local area

    Addison Group

    Houston, TX
    21 days ago
  • $162.6k - $244k

     ...Technologies, Inc. Job Area Engineering Group, Engineering...  ...Qualcomm Data Center AI System Hardware and Validation...  ...team develops rack-level AI inference solutions...  ...generation Qualcomm rack‑scale designs within a matrixed...  .... Drive firmware/software/hardware co‑design execution... 
    Work experience placement
    Work at office
    Work from home

    Qualcomm

    Austin, TX
    2 days ago
  •  ...Quicknode is a cloud-based infrastructure company that powers the blockchain...  ...As a Senior Infrastructure Engineer at Quicknode, you’ll play a...  ...on building robust, scalable systems that go beyond traditional cloud...  ...diagrams to support scale, reliability, and compliance... 
    Full time
    Remote work
    Flexible hours

    QuickNode

    Remote
    2 days ago
  • $61k - $101k

     ...mechanical or electrical engineering. We need strong...  ...understanding of production rack environments across...  ...AI tools to support infrastructure engineering workflows,...  ...them into end-to-end system designs. We drive continuous...  ...across hardware and software stakeholders. We... 
    Full time

    J.P. Morgan

    Houston, TX
    4 days ago
  • $182k - $242k

     ...enables innovators to build and scale AI with confidence. Trusted by...  ..., CoreWeave combines superior infrastructure performance with deep...  ...About the Team: The IT Engineering at CoreWeave designs, builds, and operates the systems that enable our employees and... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    14 days ago
  •  ...scientific programs. We build the infrastructure behind some of the most...  ...seeking a talented and versatile Software / API Engineer to join a dynamic team building software systems that integrate scientific workflows and large-scale supercomputing environments. This... 

    LTD Global

    Berkeley, CA
    26 days ago
  • Lead Infrastructure Engineer Assume a vital position as a key member of a high...  ...you apply deep knowledge of software, applications, and...  ...upstream and downstream data and systems or technical implications and...  ..., including cluster setup, scaling, and CI/CD pipeline integration... 
    Shift work

    Hackajob

    Plano, TX
    3 days ago
  • $125.84k - $238.16k

     ...The Senior Staff Software Engineer manages assigned software systems through strategic direction,...  ...operational health. The Principal Software Engineer...  ...facilitate genomics analysis at scale. Foundationally, this...  ..., Rust Genomics Infrastructure. Explore our exceptional... 
    Full time
    Local area

    St Jude Children's Research Hospital

    Memphis, TN
    a month ago
  •  ...Job Description Job Description Senior Software Engineer – Backend Systems & Infrastructure Location: San Francisco - on-site Compensation: Competitive...  ...autonomous robotic platforms operating at scale.   The ideal candidate enjoys solving complex engineering... 

    MRINetwork Jobs

    San Francisco, CA
    23 days ago
  • $170k - $250k

     ...CPAs, tax attorneys, and engineers, Taxbit is the leading innovator...  ...is looking for a Staff Software Engineer to join our Systems, Security & Compliance...  ...ownership of the cloud infrastructure that powers our platform....  ...systems responsibly and at scale. You will have broad... 
    Work at office
    Work from home
    Flexible hours

    TaxBit

    Washington DC
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!