Principal Software Engineer - Rack Scale Systems Infrastructure
$272k - $431.25kDormont Manufacturing Co
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.
What You Will Be Doing:
Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.
Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.
Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.
Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.
Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.
Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long‑term infrastructure evolution.
What We Need To See:
BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.
Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.
Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production‑quality infrastructure software. Experience with Rust is highly valued.
Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large‑scale infrastructure services. Experience with Linux‑based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring‑up workflows.
Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric‑level manageability. Experience with complex accelerator‑based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high‑performance computing systems.
Expertise in in‑band and out‑of‑band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.
Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment‑specific integrations.
Experience using AI‑assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.
Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.
Ways To Stand Out from the crowd:
Built software supporting multiple adoption models — internal services, CSP‑integrated offerings, reusable libraries, and customer‑extensible APIs. Strong Rust skills in systems, infrastructure, or hardware‑adjacent software.
Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI‑assisted engineering. Hands‑on with fleet‑scale provisioning, updates, rollback, observability, health, and remediation.
Led across the full data center product lifecycle: inception, pre‑and post‑silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.
Deep experience with rack‑or cluster‑scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long‑term direction.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people in the world working for us. If you’re creative and autonomous, we want to hear from you!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until May 19, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
#J-18808-Ljbffr$272k - $431.25k
...productivity required for strong scaling for HPC and generative AI... ...are looking for expert engineers to come and help design rack level solutions for next... ...solutions for scaling AI infrastructure using GPUs and Grace... ...space complexity and project system resource requirements....Suggested$272k - $431.25k
...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations in automotive platform software, system... ...to the architecture, development, optimization, and scaling of foundational software technologies powering NVIDIA automotive...Suggested$120k - $150k
A mission-driven nonprofit is hiring a Cloud Infrastructure Engineer to manage and optimize their Google Cloud Platform. Ideal candidates will... ...services, particularly Google Cloud. Responsibilities include system administration, and data management to ensure high platform...SuggestedRemote job$100k
Dormont Manufacturing Co is looking for an AI Scale-Out Software Engineer to build and optimize our TT-fabric and distributed training infrastructure. This position requires expertise in deep learning and low-level networking. The ideal candidate will work collaboratively...Suggested$272k - $431.25k
...We are hiring senior engineers to work on the CUDA driver, a core component of our platform... ...programming model across a range of system configurations and hardware capabilities... ...experience) ~15+ years of relevant systems software development experience ~ Strong C...Suggested- ...generation of computer user agents - AI systems that can actually use your computer... ...We’re looking for a generalist backend/infrastructure engineer who thrives in ambiguity, has strong architectural... ...- you’ll be shaping how our platform scales and evolves over the next few years....
$136k - $218.5k
...customer-facing hardware engineers to work directly with Cloud Scale Providers (CSP’s)... ...Centers. The HW Systems Engineer is front-... ...Vera Ruben NVL72 racks, at our largest... ...compute and networking infrastructure needed for agentic... ...at the hardware, software and application...$184k - $287.5k
...built. We are seeking a Senior Software Engineer focused on container and cloud infrastructure. You will help design and implement... ...reliability, performance, and scale across thousands of GPUs. There... ...Expertise with Helm chart design systems, Operators, and platform APIs serving...$272k - $431.25k
...personal assistants and engineering‑productivity... .... Now we need a principal‑level, hands‑on... ...generation of agent infrastructure. This is not a... ...harden production systems and the architectural... ...to ensure they scale. What you’ll be... ...like mature software, not prototypes....Live in$130k - $180k
Constellation Software Engineer, Infrastructure WindBorne Systems is supercharging weather forecasts with a unique proprietary data source: a global constellation... ...a Constellation Infrastructure Software Engineer to scale and build up the mission critical systems that drive...Work at officeImmediate start- ...more. What You’ll Do As a Principal Java Software Engineer, you’ll set the technical... ...highly scalable, distributed systems in network security. Lead... ...-generation capabilities. Scale the platform for 10x+... ...on experience with network infrastructure, including routers, firewalls...Full timeRemote workWorldwideFlexible hours
- ...Hash is looking for a Staff Software Engineer to architect and build our... ...data aggregation platform, scaling our gateway to Finance 2.0.... ...the challenges of complex systems and browser internals. Because... ...role is critical to our infrastructure. You will be responsible for...Work experience placementWork from home
$190k - $270k
Staff Software Engineer - AI Research Infrastructure P-1215 At the company, we are obsessed with enabling data teams... ..., orchestrate, and observe large‑scale training and inference experiment... ...clusters, GPU fleets, or cloud‑based systems) Enable researchers to go from...Local area$197.3k - $313.7k
Software Engineering About Salesforce Salesforce is the #1 AI CRM, where humans... ...within the Architecture & Systems organization. While most... ...existing frontend and desktop infrastructure to support new product... ...modern web applications at scale. Deep Chromium experience,...$184k - $287.5k
...are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect... ...and extend high-level DSLs and compiler infrastructure to boost kernel developer productivity...- ...technology. We are at the forefront of software and hardware innovation,... ...that enables the software infrastructure to scale and to support the entire software engineering organization. As part of the software... ...of our ML accelerator systems both on hardware and software....Work experience placementRemote work3 days per week
$170k - $220k
.... Our platform provides engineers with real-time observability... ..., and Starship—where scaling telemetry, debugging flight systems, and ensuring mission reliability demanded new infrastructure. Founded by a team from... ...negotiable. About the Role As a Software Engineer, Security...Permanent employmentWork at officeRelocation$232k - $290k
...idea to global enterprises scaling their digital presence, we... ...looking for a Senior Staff Software Engineer to help lead the technical... ...complex, revenue-generating system and figuring out what it needs... ...investments and the data infrastructure that supports them, that unlock...Permanent employmentFull timeTemporary workFixed term contractImmediate startRemote workFlexible hours$120k - $150k
OpenReview Is Growing — We’re Hiring a Cloud Infrastructure Engineer We're cross posting for our colleagues at OpenReview As OpenReview... ...monthly users and continuous year-over-year growth, we’re scaling our systems and our team to meet the demands of this vibrant...Remote work- Seer is seeking a Systems Engineer to design and operate infrastructure for large-scale AI systems. You'll build solutions for distributed training, ensure reliability, and work with advanced GPU clusters. Candidates should have over 3 years of relevant experience and deep...
$20k
...’re looking for someone to lead all technical aspects of an engineering team at ServiceTitan. You must have a strong background in responsive web application development, building distributed systems for scale, and a proven ability to deliver technical leadership and strong...Minimum wageTemporary workLocal areaFlexible hours$165.8k - $307.9k
...Role Overview As a Principal Software Developer in Test, you will be responsible... ...you will represent quality engineering and verification on behalf... ...testing level (end‑to‑end, system, integration, unit, etc.)... ...methodologies including Scrum and Scaled Agile Framework; databases...Work at officeLocal areaRelocation package$272k - $431.25k
...deeply technical, hands‑on Principal Engineer to lead the security foundations... ...observability and auditing infrastructure (structured logs, decision... ...and operationally safe at scale Design and operate a... ...building and securing large‑scale systems, platforms, or...$184k - $287.5k
...means contributing to the infrastructure that powers our innovative... ...the necessary resources and scale to foster innovation. We are... ...seeking an AI infrastructure software engineer to join our team. You’ll be... ...implementing software and systems engineering practices to...$168k - $270.25k
...NVIDIA, Site Reliability Engineering provides a rare chance to... ...develop, and support large-scale production systems with high efficiency and availability... ...demanding position merges software and systems engineering... .... Background with infrastructure automation. Experience running...$100k - $137k
# DevOps Engineer - ML & Data InfrastructureHigh 5 GamesFull TimejuniorCAPosted... ...DevOps Engineer - ML & Data Infrastructure. This is a full-time role in... .... You’ll play a key role in scaling AI models from research to... ..., and keep our AI systems running seamlessly for...Full timeWorldwide$200.6k - $250.4k
...‑class technical leadership—engineers who can see across domains, design foundational systems, and set the architectural direction... ...for years to come. As a Principal Staff Software Engineer, you will play a... ...platform problems at global scale, influencing enterprise...Flexible hours- Role Overview As a Principal Staff Engineer at Jazzx.ai, you will play a pivotal... ...of building scalable systems, deep experience building... ...AI platform services and infrastructure. Drive the evolution of our... ...reliable, and user-centric software solutions. Lead architectural...
$184k - $287.5k
...doing: Develop use cases and system requirements for L3 and L4... ...closely with Data Analytics, Test Engineering, and System Integration &... ...analysis, data analysis, and software architecture. Strong software... ...Hands‑on experience with large-scale datasets, data science, and...$180k - $280k
...builds autonomous aerospace systems for the real world. At the heart... ...real time, safety critical software stack that has to work under... ...a Senior Embedded Software Engineer, you will own software that... ..., CMake, custom build infrastructure Version control and CI: Git...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!
- principal software engineer California, MO
- systems engineer California, MO
- advanced systems engineer California, MO
- space systems engineer California, MO
- senior linux systems engineer California, MO
- mission system engineer California, MO
- operating system engineer California, MO
- software system engineer California, MO
- distributed systems engineer California, MO
- senior staff systems engineer California, MO


