Principal Software Engineer - Rack Scale Systems Infrastructure
$272k - $431.25kNVIDIA
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.What You Will Be Doing:Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.What We Need To See:BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.Ways To Stand Out from the crowd:Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US, Remote; US, MA, RemoteType: Full time
$120k - $165k
...saves livesWe're seeking a Principal Software Engineer to lead the architecture... ...audio and video communication infrastructure powering our Connected... ...you'll own mission-critical systems that enable seamless, real... ...-time communications at scale.Scale and optimize media servers...SuggestedFull timeTemporary workWork visaFlexible hours3 days per week$272k - $431.25k
...NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. We expect you to have significant software engineering experience... ...DGX Cloud team responsible for production systems that enable large scalable GPU clusters to...SuggestedFull time$184k - $287.5k
...the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely... ...performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies...SuggestedFull timeRemote work$152k - $241.5k
...a team that analyzes large-scale datacenter workloads on GPU... ...with OS, container, GPU, and systems engineers. When useful, you will... ...large-scale workloads and infrastructure signals to find application... ...prediction) inside existing software workflows.What we need to see...SuggestedFull timeRemote work- ...companies and health systems make confident decisions... .... With our global scale and deep expertise, you... ...of experience in software engineering.Preferred Qualifications... ...with Terraform for infrastructure as code and AWS resource... ..., audits).The Principal Software Engineer is...SuggestedFull timeTemporary workCasual workInternshipMonday to FridayFlexible hoursDay shift
$231.4k - $331.8k
...Isovalent builds open-source software and enterprise... ...modern cloud native infrastructure. The flagship technologies... ...environments. As a Principal Engineer, you will be at the... ...in building the systems that power our platform... ...things happen on a global scale. Because our...Full timeTemporary workLocal areaFlexible hours$222.9k - $295.2k
...observability platform to keep their systems secure, available, and... ...ImpactWe're looking for a Principal Software Engineer who enjoys solving... ...designing and building large-scale distributed systems or... ...revolutionizing how data and infrastructure connect and protect...Full timeTemporary workLocal areaFlexible hours$110k - $180k
...techWe are seeking a hands-on Principal Software Engineer to lead the design and delivery of enterprise-scale business automation and AI... ..., and event-driven systems. Proven ability to establish... ...AI systems. Experience with Infrastructure as Code, API design and lifecycle...Full timeTemporary workPart timeWork experience placementWork at officeRemote workRelocation packageFlexible hours$169k - $224k
...organization of scientists, engineers, and physicians and we... ...(NGS), population-scale clinical studies, and state... ...grail.comThe Staff Software Development Engineer - Enterprise AI Infrastructure is a senior technical role... ..., and multi-agent systems using advanced LLM orchestration...Full timeTemporary workWork at officeLocal areaFlexible hoursShift work- ...the following job description:The Lead Infrastructure Engineer within Truist’s Digital Workplace... ...infrastructure technology platforms and systems across various environments, such as cloud... ...Vulnerability Management (Enterprise Scale)Operate and mature the endpoint...Permanent employmentFull timePart timeWork experience placementH1bWork at officeWork visaShift workDay shift
- ...Investment Bank Technology (CCIBT) Infrastructure Solutions organization is seeking a hands-on Lead Engineer to design and deliver... ...within highly regulated, large-scale enterprise environments. This... ...design.Experience with distributed systems, APIs, integration, and...Full timeWork experience placement
- ...implements enterprise platform software services that enable... ...event-driven automation systems, and policy-as-code enforcement engines that integrate self-... ...reliability at enterprise scale. Applies modern software... ...implements enterprise infrastructure technology platforms and...Full timePart timeShift workDay shift
- ...strategic Observability Lead-Sr. Infrastructure Engineer with a strong Forward... ...focus to define, implement, and scale enterprise observability... ...embed observability into the software development lifecycle, and ensure... ...position means improving system reliability, reducing mean-...Permanent employmentFull timePart timeWork experience placementH1bWork at officeWork visaShift workDay shift
$114.6k - $234.6k
...Broomfield, CO Overview Oracle Cloud Infrastructure (OCI) is building Oracle Video @... ...a highly technical, distributed systems-focused engineering team Responsibilities Responsibilities... ...~ Experience building high-scale, low-latency systems ~ Bachelor's...Temporary workFlexible hours- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure... ...accelerated workloads, large-scale models, simulations, and... ...solutions, aligning system architecture and deployment roadmaps... ...along with LLM inference engines (TensorRT-LLM), production serving...Full timeWork experience placementLive inWork at officeLocal area
$100k - $160k
...Our AI Voice & Telephony Engineering Team builds the systems that power intelligent conversations... ...-like interactions at scale. We’re driving the next... ...’ll Do: As a Senior Software Engineer, you’ll design... .../Node.js. Cloud Infrastructure: Deep experience with AWS...Full timeTemporary workFlexible hours$147k - $294k
...appliance that can incrementally scale out to manage petabytes of... ...we are looking for a Software Engineer in the Core Data Path that are... ...shared-nothing distributed file system) is the foundational piece... ...defined storage helps us power infrastructure for all kinds of...Work at officeRemote workRelocation package3 days per week- ...Summary Our Deloitte AI & Engineering team to transform technology... ...on client requirementsLead system migrations and upgrades to... ...promise. Our Hybrid Cloud Infrastructure offering provides specialized... .... It leverages Deloitte’s scale and talent, as well as a center...Work at officeLocal area
- ...DescriptionIn the assigned Job Role of Infrastructure Consultant 2, your Area Of Responsibility... ...develop technical documentation for deployed systems Align release schedules and environment... ...the business with agile digital at scale to deliver unprecedented levels of performance...Full timeTemporary workRemote workRelocation
$116.2k - $229.1k
...Join Deloitte’s AI & Engineering practice and help organizations... ...Azure. As an Azure/Databricks Infrastructure Engineer, you will design, implement... ...cloud operations, and scale secure, intelligent platforms... ..., Engineering, Information Systems, Analytics, or related field...Local areaVisa sponsorship$134k - $265k
...Domain Tech Lead — Infrastructure Patching Automation (Mainframe... ...AS/400)Join our AI & Engineering team by transforming... ...sector solutions in software, data, AI, network, and... ...domain on a large-scale infrastructure engineering... ...up to 50%IBM z/OS systems programming and patching...Local area$207k - $301k
...future requirements and infrastructure needs.Design, guide and vet systems designs within the scope... ...code developed by other engineers and provide feedback to... ...years of experience in software development.3 years of experience... ...with developing large-scale infrastructure,...- ...initiatives and deliverables within Software Engineering and contribute to large-scale planning related to Software... ...experience, education. Experienced Infrastructure Engineering professional with 7+... ...expertise in IBM MQ ecosystems, systems administration, automation,...Work experience placement
$97.5k - $209.5k
...experience Extensive hands-on software design and programming... ...and building distributed systems and large-scale architectures Strong... ...technical guidance to junior engineers and peers Drive... ...brings together the data, infrastructure, applications, and expertise...Temporary workFlexible hours$207k - $301k
...Software Engineer Manager, Google Meet Infrastructure Note: By applying to this position you will have an opportunity to share your preferred... ...years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute...- ...role:Wells Fargo is seeking a Lead Systems Operations Engineer in technology as part of Commercial... ...tested, released, and stabilized at scale.This role embeds Site Reliability Engineering... ...of computer systems and network infrastructure for Systems Operations functional...Full timeWork experience placement
$250.6k - $362.6k
...security outcomes, as a Principal Engineer. The team delivers... ...spans cloud and client infrastructure, security, policy and... ...develop, and drive unique software and technology... ...and developing large-scale software, platform, or distributed systems.Experience building and...Full timeTemporary workLocal areaRemote workFlexible hours$122k - $200k
...for defining and leading the engineering approach for complex... ...have built through multiple software implementations and have developed... ...(this is based on the scale of implementation and skillsets... ...of experience in platform, systems, infrastructure, or cloud engineering, with...Full timeWork at officeDay shift- Lead Software Engineer — Enterprise Video Services, Wello, and Brand & Sponsorship... ...and distribution at scale.In this role, you will:Lead... ...provision of high-level systems consultation for the technology... ...systems and network infrastructure for Systems Operations functional...Full timeWork experience placementWork at officeRelocation3 days per week
- ...-hire Network Cloud Infrastructure Automation Engineer, Burlington & Ayer,... ...digital production systems. They continuously seek... .... Much of our software development focuses... ...complex challenges of scale which are unique to... ...by· Dom Costagliola, Principal, m (***) ***-****, domenic...Permanent employmentContract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!
- systems engineer North Carolina
- healthcare systems engineer North Carolina
- senior windows systems engineer North Carolina
- operating system engineer North Carolina
- senior infrastructure engineer North Carolina
- infrastructure engineer North Carolina
- security infrastructure engineer North Carolina
- infrastructure developer North Carolina
- remote infrastructure engineer North Carolina
- principal North Carolina



