Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure

$272k - $431.25k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.What You Will Be Doing:Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.What We Need To See:BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.Ways To Stand Out from the crowd:Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US, Remote; US, MA, RemoteType: Full time

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure in North Carolina vacancy
  • $120k - $165k

     ...saves livesWe're seeking a Principal Software Engineer to lead the architecture...  ...audio and video communication infrastructure powering our Connected...  ...you'll own mission-critical systems that enable seamless, real...  ...-time communications at scale.Scale and optimize media servers... 
    Suggested
    Full time
    Temporary work
    Work visa
    Flexible hours
    3 days per week

    Baxter International

    Raleigh, NC
    1 day ago
  • $272k - $431.25k

     ...NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. We expect you to have significant software engineering experience...  ...DGX Cloud team responsible for production systems that enable large scalable GPU clusters to... 
    Suggested
    Full time

    Nvidia

    Durham, NC
    13 hours ago
  • $184k - $287.5k

     ...the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely...  ...performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies... 
    Suggested
    Full time
    Remote work

    Nvidia

    North Carolina
    3 days ago
  • $152k - $241.5k

     ...a team that analyzes large-scale datacenter workloads on GPU...  ...with OS, container, GPU, and systems engineers. When useful, you will...  ...large-scale workloads and infrastructure signals to find application...  ...prediction) inside existing software workflows.What we need to see... 
    Suggested
    Full time
    Remote work

    Nvidia

    North Carolina
    3 days ago
  •  ...companies and health systems make confident decisions...  .... With our global scale and deep expertise, you...  ...of experience in software engineering.Preferred Qualifications...  ...with Terraform for infrastructure as code and AWS resource...  ..., audits).The Principal Software Engineer is... 
    Suggested
    Full time
    Temporary work
    Casual work
    Internship
    Monday to Friday
    Flexible hours
    Day shift

    Laboratory Corporation of America

    Durham, NC
    4 days ago
  • $231.4k - $331.8k

     ...Isovalent builds open-source software and enterprise...  ...modern cloud native infrastructure. The flagship technologies...  ...environments. As a Principal Engineer, you will be at the...  ...in building the systems that power our platform...  ...things happen on a global scale. Because our... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Research Triangle Park, NC
    13 hours ago
  • $222.9k - $295.2k

     ...observability platform to keep their systems secure, available, and...  ...ImpactWe're looking for a Principal Software Engineer who enjoys solving...  ...designing and building large-scale distributed systems or...  ...revolutionizing how data and infrastructure connect and protect... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Research Triangle Park, NC
    1 day ago
  • $110k - $180k

     ...techWe are seeking a hands-on Principal Software Engineer to lead the design and delivery of enterprise-scale business automation and AI...  ..., and event-driven systems. Proven ability to establish...  ...AI systems. Experience with Infrastructure as Code, API design and lifecycle... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Work at office
    Remote work
    Relocation package
    Flexible hours

    Ally Financial

    North Carolina
    13 hours ago
  • $169k - $224k

     ...organization of scientists, engineers, and physicians and we...  ...(NGS), population-scale clinical studies, and state...  ...grail.comThe Staff Software Development Engineer - Enterprise AI Infrastructure is a senior technical role...  ..., and multi-agent systems using advanced LLM orchestration... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    Shift work

    GRAIL

    Durham, NC
    11 hours ago
  •  ...the following job description:The Lead Infrastructure Engineer within Truist’s Digital Workplace...  ...infrastructure technology platforms and systems across various environments, such as cloud...  ...Vulnerability Management (Enterprise Scale)Operate and mature the endpoint... 
    Permanent employment
    Full time
    Part time
    Work experience placement
    H1b
    Work at office
    Work visa
    Shift work
    Day shift

    Truist

    Raleigh, NC
    4 days ago
  •  ...Investment Bank Technology (CCIBT) Infrastructure Solutions organization is seeking a hands-on Lead Engineer to design and deliver...  ...within highly regulated, large-scale enterprise environments. This...  ...design.Experience with distributed systems, APIs, integration, and... 
    Full time
    Work experience placement

    Wells Fargo

    Charlotte, NC
    2 days ago
  •  ...implements enterprise platform software services that enable...  ...event-driven automation systems, and policy-as-code enforcement engines that integrate self-...  ...reliability at enterprise scale. Applies modern software...  ...implements enterprise infrastructure technology platforms and... 
    Full time
    Part time
    Shift work
    Day shift

    Truist

    Charlotte, NC
    4 days ago
  •  ...strategic Observability Lead-Sr. Infrastructure Engineer with a strong Forward...  ...focus to define, implement, and scale enterprise observability...  ...embed observability into the software development lifecycle, and ensure...  ...position means improving system reliability, reducing mean-... 
    Permanent employment
    Full time
    Part time
    Work experience placement
    H1b
    Work at office
    Work visa
    Shift work
    Day shift

    Truist

    Charlotte, NC
    4 days ago
  • $114.6k - $234.6k

     ...Broomfield, CO Overview Oracle Cloud Infrastructure (OCI) is building Oracle Video @...  ...a highly technical, distributed systems-focused engineering team Responsibilities Responsibilities...  ...~ Experience building high-scale, low-latency systems ~ Bachelor's... 
    Temporary work
    Flexible hours

    Oracle

    Raleigh, NC
    13 hours ago
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure...  ...accelerated workloads, large-scale models, simulations, and...  ...solutions, aligning system architecture and deployment roadmaps...  ...along with LLM inference engines (TensorRT-LLM), production serving... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Charlotte, NC
    5 days ago
  • $100k - $160k

     ...Our AI Voice & Telephony Engineering Team builds the systems that power intelligent conversations...  ...-like interactions at scale. We’re driving the next...  ...’ll Do: As a Senior Software Engineer, you’ll design...  .../Node.js. Cloud Infrastructure: Deep experience with AWS... 
    Full time
    Temporary work
    Flexible hours

    Red Ventures

    Charlotte, NC
    13 hours ago
  • $147k - $294k

     ...appliance that can incrementally scale out to manage petabytes of...  ...we are looking for a Software Engineer in the Core Data Path that are...  ...shared-nothing distributed file system) is the foundational piece...  ...defined storage helps us power infrastructure for all kinds of... 
    Work at office
    Remote work
    Relocation package
    3 days per week

    Nutanix

    Durham, NC
    2 days ago
  •  ...Summary Our Deloitte AI & Engineering team to transform technology...  ...on client requirementsLead system migrations and upgrades to...  ...promise. Our Hybrid Cloud Infrastructure offering provides specialized...  .... It leverages Deloitte’s scale and talent, as well as a center... 
    Work at office
    Local area

    Deloitte

    Charlotte, NC
    13 hours ago
  •  ...DescriptionIn the assigned Job Role of Infrastructure Consultant 2, your Area Of Responsibility...  ...develop technical documentation for deployed systems Align release schedules and environment...  ...the business with agile digital at scale to deliver unprecedented levels of performance... 
    Full time
    Temporary work
    Remote work
    Relocation

    Infosys Technologies

    Charlotte, NC
    1 day ago
  • $116.2k - $229.1k

     ...Join Deloitte’s AI & Engineering practice and help organizations...  ...Azure. As an Azure/Databricks Infrastructure Engineer, you will design, implement...  ...cloud operations, and scale secure, intelligent platforms...  ..., Engineering, Information Systems, Analytics, or related field... 
    Local area
    Visa sponsorship

    Deloitte

    Charlotte, NC
    4 days ago
  • $134k - $265k

     ...Domain Tech Lead — Infrastructure Patching Automation (Mainframe...  ...AS/400)Join our AI & Engineering team by transforming...  ...sector solutions in software, data, AI, network, and...  ...domain on a large-scale infrastructure engineering...  ...up to 50%IBM z/OS systems programming and patching... 
    Local area

    Deloitte

    Charlotte, NC
    2 days ago
  • $207k - $301k

     ...future requirements and infrastructure needs.Design, guide and vet systems designs within the scope...  ...code developed by other engineers and provide feedback to...  ...years of experience in software development.3 years of experience...  ...with developing large-scale infrastructure,... 

    Google

    Raleigh, NC
    1 day ago
  •  ...initiatives and deliverables within Software Engineering and contribute to large-scale planning related to Software...  ...experience, education. Experienced Infrastructure Engineering professional with 7+...  ...expertise in IBM MQ ecosystems, systems administration, automation,... 
    Work experience placement

    Mindlance

    Charlotte, NC
    3 days ago
  • $97.5k - $209.5k

     ...experience Extensive hands-on software design and programming...  ...and building distributed systems and large-scale architectures Strong...  ...technical guidance to junior engineers and peers Drive...  ...brings together the data, infrastructure, applications, and expertise... 
    Temporary work
    Flexible hours

    Oracle

    Raleigh, NC
    3 days ago
  • $207k - $301k

     ...Software Engineer Manager, Google Meet Infrastructure Note: By applying to this position you will have an opportunity to share your preferred...  ...years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute... 

    Habitat For Humanity Of Durham

    Raleigh, NC
    1 day ago
  •  ...role:Wells Fargo is seeking a Lead Systems Operations Engineer in technology as part of Commercial...  ...tested, released, and stabilized at scale.This role embeds Site Reliability Engineering...  ...of computer systems and network infrastructure for Systems Operations functional... 
    Full time
    Work experience placement

    Wells Fargo

    Charlotte, NC
    4 days ago
  • $250.6k - $362.6k

     ...security outcomes, as a Principal Engineer. The team delivers...  ...spans cloud and client infrastructure, security, policy and...  ...develop, and drive unique software and technology...  ...and developing large-scale software, platform, or distributed systems.Experience building and... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Research Triangle Park, NC
    4 days ago
  • $122k - $200k

     ...for defining and leading the engineering approach for complex...  ...have built through multiple software implementations and have developed...  ...(this is based on the scale of implementation and skillsets...  ...of experience in platform, systems, infrastructure, or cloud engineering, with... 
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, NC
    4 days ago
  • Lead Software Engineer — Enterprise Video Services, Wello, and Brand & Sponsorship...  ...and distribution at scale.In this role, you will:Lead...  ...provision of high-level systems consultation for the technology...  ...systems and network infrastructure for Systems Operations functional... 
    Full time
    Work experience placement
    Work at office
    Relocation
    3 days per week

    Wells Fargo

    Charlotte, NC
    3 days ago
  •  ...-hire Network Cloud Infrastructure Automation Engineer, Burlington & Ayer,...  ...digital production systems. They continuously seek...  .... Much of our software development focuses...  ...complex challenges of scale which are unique to...  ...by· Dom Costagliola, Principal, m (***) ***-****, domenic... 
    Permanent employment
    Contract work

    BlueSkyClarity

    Burlington, NC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!