Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure

$272k - $431.25k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.What You Will Be Doing:Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.What We Need To See:BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.Ways To Stand Out from the crowd:Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US, Remote; US, MA, RemoteType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure in Raleigh, NC vacancy
  • $114.6k - $234.6k

     ...Description Owns complex system modules and fleet...  ...integration contracts across software, firmware, and...  ...provisioning pipelines at scale. Elevates code quality...  ...and coaching to engineers to drive improvements....  ...brings together the data, infrastructure, applications, and expertise... 
    Suggested
    Temporary work
    Flexible hours
    Shift work

    Oracle

    Raleigh, NC
    3 days ago
  • $114.6k - $234.6k

     ...We are looking for smart systems software engineers with BS/MS/PhD in Computer...  ...X11, and Exadata Expansion Rack. Exadata group also keeps looking...  ...distributed algorithms to scale systems Career Level - IC4...  ...brings together the data, infrastructure, applications, and... 
    Suggested
    Temporary work
    Flexible hours

    Oracle

    Raleigh, NC
    3 days ago
  • $184k - $287.5k

     ...the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely...  ...performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies... 
    Suggested
    Full time
    Remote work

    Nvidia

    Raleigh, NC
    4 days ago
  • $152k - $241.5k

     ...a team that analyzes large-scale datacenter workloads on GPU...  ...with OS, container, GPU, and systems engineers. When useful, you will...  ...large-scale workloads and infrastructure signals to find application...  ...prediction) inside existing software workflows.What we need to see... 
    Suggested
    Full time
    Remote work

    Nvidia

    Raleigh, NC
    4 days ago
  • $143k - $193.05k

     ...Description Summary: The Infrastructure Modernization business unit is seeking a Senior Principal Software Engineer to serve as a visionary and...  ...in z Assembler (HLASM), system diagnostics, and the z/OS architecture...  ...and roadmap for enterprise-scale solutions. ~ Experience... 
    Suggested
    Full time
    Remote work
    Worldwide

    Rocket Software

    Raleigh, NC
    2 days ago
  • $117.1k - $187.3k

     ...Lead all aspects of critical systems for delivering innovative...  ...will build and lead remote software engineering organizations using agile principles...  ...and deliver products that scale to support our millions of...  ...Understanding of Infrastructure as Code (IaC) tools (e.g.,... 
    Full time
    Work experience placement
    Live in
    Local area
    Remote work
    Worldwide

    Cengage Group

    Raleigh, NC
    4 days ago
  • $186.16k - $304.7k

    Job Title: Principal Software Engineer Work Location: 145 Broadway, Cambridge,...  ...in code reviews to ensure systems are evaluated holistically...  ...measurements. Identify when large-scale refactoring is necessary...  ...with QA on creating test infrastructure and automation. Find OS-... 
    Work experience placement
    Work at office
    Remote work

    Akamai

    Raleigh, NC
    3 days ago
  • $135.2k - $306.4k

     ...Description Oracle Cloud Infrastructure (OCI) delivers...  ...offers unmatched hyper-scale, multi-tenant services...  ...hoping to enhance engineering efficiency by concentrating...  ...building low level systems with high performance...  ...As a Senior Principal Engineer, you will lead... 
    Temporary work
    Worldwide
    Flexible hours

    Oracle

    Raleigh, NC
    3 days ago
  • $152.07k - $202.76k

     ...communities. At Lumen, you’ll work on infrastructure customers rely on today and...  ...us today. The Role The Principal Software Engineer serves as a senior technical...  ...-native engineering, agentic systems, intelligent automation, and large-scale distributed software... 
    Temporary work
    Remote work

    Lumen Inc

    Raleigh, NC
    3 days ago
  • $184k - $287.5k

     ...scheduler cells by hand does not scale, and we are not going to try....  ..., and we need an automation engineer to own it end to end. You are...  ...be migrating off partial systems and reconciling drift, not starting...  ...experience.6+ years in infrastructure engineering with strong config... 
    Full time
    Remote work

    Nvidia

    Raleigh, NC
    2 days ago
  • $272k - $431.25k

     ...working for us! The Cloud Engineering & Services team is building...  ...governance layer for agentic systems: signed policy, runtime...  ...integrations.We are looking for a Principal Software Engineer, Agent Policy...  ...Engage alongside OpenShell and Infrastructure engineers on public runtime... 
    Full time
    Contract work
    Remote work

    Nvidia

    Raleigh, NC
    5 days ago
  • $150k - $300k

     ...EngineeringCity: RaleighState: NC Veeva Systems is a mission-driven organization and...  ...employees, and communities.The RoleAs Principal Software Engineer for a new product within Veeva, you...  ...Extensive experience developing high-scale enterprise SaaS cloud applicationsScalability... 
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours
    3 days per week

    Veeva Systems

    Raleigh, NC
    1 day ago
  •  ...across multiple product and engineering teams while staying close to...  ...code, data, user experience, infrastructure, and production operations....  ...teams improve how they build software, including the responsible...  ...bottleneck. Identify and address systemic risks in security,... 
    For contractors
    Work experience placement
    Work at office
    Local area
    2 days per week

    Veradigm

    Raleigh, NC
    5 days ago
  • $184k - $287.5k

    NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on...  ...for an experienced systems software engineer who can...  ...compiler, runtime, library, infrastructure, or component owner....  ..., Perforce-scale development, or distributed... 
    Full time
    Remote work

    Nvidia

    Raleigh, NC
    4 days ago
  • $120k - $165k

     ...for patients around the world. As a Principal Full Stack Software Engineer at Baxter, your work contributes directly...  ..., and high-performing backend systems that support enterprise healthcare applications...  ...API's, and building systems at scale.Proven understanding of database... 
    Full time
    Temporary work
    Local area
    Flexible hours
    Shift work

    Baxter International

    Raleigh, NC
    4 days ago
  • $130k - $180k

     ...Qualus as a Lead Relay Settings Engineer. In this role you will...  ...interconnections to a client’s electrical system.Ensures all relay settings...  ...able to identify necessary software to complete needed design...  ...at the forefront of power infrastructure transformation, with... 
    Temporary work
    Flexible hours

    Qualus

    Raleigh, NC
    4 days ago
  • $150k - $300k

     ...Principal Software Engineer - Python Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies...  ...SaaS Leader: Extensive experience developing high-scale enterprise SaaS cloud applications Scalability... 
    Work at office
    Local area
    Work from home
    Flexible hours
    3 days per week

    Veeva Systems

    Raleigh, NC
    4 days ago
  •  ...ensuring compatibility with existing systems and enterprise processes....  ..., monitor, and optimize AI infrastructure, working with server, cloud, and platform engineering teams.Operationalize machine learning...  ...:Bachelor’s degree in CS, Software Engineering or other IT-... 
    Full time
    Work at office
    Remote work

    Applied Research Associates

    Raleigh, NC
    5 days ago
  • $168.98k - $218.68k

     ...the platform layers that let AI systems and agents find, trust, ask, and...  ...certification and governance patterns that scale data product trust from...  ...casesEngineering & InfrastructureLead engineering of scalable, secure platform infrastructure on AWS (S3, Lake Formation, Glue,... 
    Full time
    Contract work
    For contractors
    Local area

    GILEAD Sciences

    Raleigh, NC
    4 days ago
  • $142.8k - $274.8k

     ...The Microsoft Azure team is seeking a Principal Software Engineer to lead the development of firmware...  ...libraries for ANC, Compute, and SoC infrastructure subsystems. Influence...  ...C/C++ for Compute engines, ANC sub system, DMA Engines, and other similar IP blocks... 
    Ongoing contract
    Local area

    Microsoft Corporation

    Raleigh, NC
    3 days ago
  • $113k - $216.6k

    Job DescriptionJOB TITLE: Lead Engineer, Energy and Infrastructure Projects (Raleigh, NC)LOCATION: Raleigh, NC (Hybrid)Who we are?Built on more than...  ...Gas‑Insulated Switchgear (GIS), Battery Energy Storage Systems (BESS), Rail and Infrastructure programs, Data Centers,... 
    Full time
    Contract work
    Work at office

    Linxon

    Raleigh, NC
    5 days ago
  • $63k - $140k

     ...SectorNot ApplicableSpecialismPlatform Engineering & ArchitectureManagement LevelAssociateJob...  ...an Amazon Web Services, Cloud Infrastructure Experienced Associate you will support...  ...Science, Information Technology, Information Systems, Engineering- At least one of the following... 
    Full time
    Internship
    H1b

    PwC

    Raleigh, NC
    5 days ago
  • $63k - $140k

     ...& SummaryThe OpportunityAs a GenAI Python Systems Engineer - Experienced Associate, you will leverage...  ...practice, you will apply data, algorithms, and software engineering to build and deploy AI and Machine Learning solutions at scale, enabling informed decision-making and... 
    Full time
    H1b

    PwC

    Raleigh, NC
    5 days ago
  •  ...established leader in Data Center Infrastructure Management, is looking for a DevOps and Infrastructure Engineer to join our engineering team....  .... You will work closely with software engineering teams to improve...  ...processes for critical systems and applications, including local... 
    Full time
    Local area
    Remote work
    Worldwide

    Sunbird Software Inc.

    Raleigh, NC
    14 days ago
  • $65 - $70 per hour

     ...fast-growing global IT and Engineering services firm delivering innovative...  ...Engineer (SRE) – Security Infrastructure Location: Onsite 3 days...  ...excellence for a large-scale network security transformation...  ...Skills ~5+ years of SRE, systems engineering, or infrastructure... 
    Local area
    3 days per week

    MatchPoint

    Cary, NC
    3 days ago
  •  ...CVS Health is seeking a Principal Software Engineer to lead hands-on development of a real-time clinical decision-support platform on Google Cloud. You will design, implement, and ship high‑performance components, including LLM pipelines and production ML workflows, while... 

    Jobleads-US

    Raleigh, NC
    5 days ago
  • Overview Lead Infrastructure Engineer at Truist. The position description describes a lead observability engineer responsible for architecting...  ...education and work experience. In-depth knowledge in information systems and ability to identify, apply, and implement best... 
    Work experience placement
    Work at office

    Truist Inc

    Raleigh, NC
    4 days ago
  •  ...the global communications software company and Tier 1 network...  ...messaging, and emergency services infrastructure behind the brands and apps...  ...for a Senior Software Engineer who leads by example —...  ...challenge of designing and scaling distributed systems that millions of people... 
    Shift work

    Bandwidth

    Raleigh, NC
    2 days ago
  •  ...global communications software company and Tier 1 network...  ...emergency services infrastructure behind the brands and...  ...Looking For:As a Senior Engineer on our Account...  ...account provisioning system. The team enables Bandwidth...  ...at a large and growing scale. We’re looking for people... 
    Flexible hours
    Shift work

    Bandwidth

    Raleigh, NC
    2 days ago
  • $157.25k - $195.68k

     ...and orchestrating at scale within enterprise-grade...  ...implementing or maintaining systems against ISO-26262...  ...services, tools and infrastructure.Must have one (1) year...  ...RHEL composes; release engineering processes, including triaging...  ...open source software solutions, using a community... 
    Contract work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Red Hat

    Raleigh, NC
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!