Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure (Santa Clara)

$272k - $431.25k
Part-time

Nvidia

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.What You Will Be Doing:Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.What We Need To See:BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.Ways To Stand Out from the crowd:Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 21, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US, Remote; US, MA, RemoteType: Full time

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure (Santa Clara) in Santa Clara, CA vacancy
  • $272k - $431.25k

     ...We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure...  ...protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Remote; US,... 
    Suggested
    Full time
    Part time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $208k - $333.5k

     ...Architect to work in IPP's (Infrastructure, Planning and...  ...of NVIDIA's software engineers worldwide. The cloud...  ...with various operating systems (Windows/Linux/Android...  ...requirements primarily Rack Scale AI Products.Finding...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,... 
    Suggested
    Full time
    Part time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $272k - $431.25k

     ...accelerators feel like a single system at datacenter scale. As large language...  ....We are seeking a Principal Systems Engineer to define the vision and...  ...storage, or ML systems infrastructure in C/C++ and Python, with...  ...: US, CA, Santa Clara; US, WA, Remote; US, MA... 
    Suggested
    Full time
    Part time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $248k - $391k

     ...highly skilled Principal Software Engineer to join our dynamic...  ...of our infrastructure both on-prem and...  ...class AI inference systems. Join us in this...  ...inference platform scaling to frontier-class...  ...for pre-release, rack-scale GPU...  ...SummaryLocation: US, CA, Santa Clara; US, RemoteType:... 
    Suggested
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    7 hours ago
  •  ...gaming and embedded systems. Grounded in a culture...  ...Customer Program Manager - Rack Scale CPU SolutionsTHE...  ...business units and engineering organizations while driving...  ..., silicon vendors, infrastructure partners, and direct...  ...:Austin, TXSanta Clara, CAThis role is not eligible... 
    Suggested
    Part time

    AMD

    Santa Clara, CA
    2 hours ago
  • $272k - $431.25k

     ...We're looking for a Principal Software Engineer to join our CSP Engagements...  ...GPU firmware and GPU system software, working...  ...firmware at fleet scale. You will drive work...  ...across multiple GPUs in a rack-scale...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,... 
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

     ...looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is...  ...practices in large-scale GPU infrastructure, delivering powerful...  ...SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX,... 
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...edge hardware and software innovation to deliver...  ...forward‑thinking engineers tackling some of...  ...searching for a Senior Systems Software Engineer...  ...problems at large scale and help shape how AI infrastructure runs in production...  ...: US, CA, Santa Clara; US, WA, SeattleType... 
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $208k - $260k

     ...manage their hybrid cloud infrastructure. Gigamon has served...  ...organizations.As a Principal Software Engineer on the Network Management System team, you will lead the...  ...role is based out of our Santa Clara, CA headquarters,...  ...platforms that support large-scale deployments and long-... 
    Part time
    Local area
    Worldwide
    3 days per week

    Gigamon

    Santa Clara, CA
    2 hours ago
  • $320k

     ...single node HGX/DGX systems all the way up to...  ...NVLink domain rack architectures. These...  ...NVIDIA AI and HPC software stack. We're searching...  ...to drive the engineering roadmap and innovation...  ...for NVIDIA's rack-scale productsMaintain...  ...SummaryLocation: US, CA, Santa Clara; US, RemoteType:... 
    Full time
    Part time
    Shift work

    Nvidia

    Santa Clara, CA
    7 hours ago
  •  ...at the forefront of software and hardware innovation...  ...the company's server systems and rack-level product lines....  ...L10–L12)• Design and scale the L10–L12 partner network...  ...Chain, Business, Engineering, or related; 5+ years...  ...hyperscaler or AI/ML infrastructure customer dynamics;... 
    Part time

    d-Matrix

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...is seeking a Senior System Architect:...  ...Failure Attribution at Scale. As EDA or equivalent...  ...waste. We need an engineer to develop and build...  ...between hardware faults, infrastructure instability, and software defects.What you'll...  ...: US, CA, Santa Clara; US, MA, Westford;... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...edge hardware and software innovation to...  ...team of innovative engineers dedicated to solving...  ...Senior Systems Software Engineer...  ...world problems at scale. In this pivotal...  ...challenge of scaling AI infrastructure while optimizing...  ...SummaryLocation: US, CA, Santa Clara; US, WA,... 
    Full time
    Part time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 hours ago
  •  ...forefront of software and hardware innovation...  ...onsite at our Santa Clara, CA,...  ...PrincipalSystem Software Engineer - AI Inference...  ...to build and scale software...  ...a team of system software experts...  ...the deployment infrastructure, working...  ...EngineeringCompensationIC6 Principal$195K – $285K •... 
    Part time
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 hours ago
  • $272k - $431.25k

     ...organization seeks a Principal Engineer to architect and scale next-generation L1...  ...L11 diagnostic systems for Cloud Service...  ...systems and hardware / software interfaces is...  ...unique data center infrastructures.What we need to see...  ...: US, CA, Santa Clara; US, CA, RemoteType... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $272k - $431.25k

     ...re looking for a Principal Engineer to join our CSP...  ...patterns and drive systemic improvements in...  ...latest NVIDIA rack-scale systems, GPU architectures...  ...configuration, software, or workload...  ...work (profiling infrastructure, benchmark...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin;... 
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $272k - $431.25k

     ...We're looking for a Principal Software Engineer to join our CSP Engagements...  ...point for fleet-scale reliability, working...  ...enables you to distinguish systemic architectural gaps...  ...in multi-NUMA, rack-scale system software...  ...SummaryLocation: US, CA, Santa ClaraType: Full time... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU...  ...decisions surrounding topics such as infrastructure, continuous integration and...  ...#deeplearningSummaryLocation: US, CA, Santa Clara; US, IL, ChampaignType: Full time... 
    Full time
    Part time
    Work experience placement

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

     ...world.Join NVIDIA's software infrastructure team to design, build...  ...improve software systems for rack, networking, and datacenter...  ...a Senior Software Engineer - Datacenter Systems...  ...supporting large-scale GPU clusters connected...  ...: US, CA, Santa Clara; US, CA, Remote; US,... 
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $272k - $431.25k

     ...required for strong scaling for HPC and...  ...looking for expert engineers to come and help design rack level solutions for...  ...solutions for scaling AI infrastructure using GPUs and...  ...complexity and project system resource...  ...SummaryLocation: US, CA, Santa Clara; US, RemoteType:... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

     ...We are looking for a senior systems software engineer to advance the Continuous...  ...Integration and Delivery (CI/CD) infrastructure that powers many...  ...operating centralized and large-scale automated build, test, and...  ...by law.SummaryLocation: US, CA, Santa ClaraType: Full time... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $208k - $327.75k

     ...seeking a Senior System Architect to define...  ..., and hands-on infrastructure validation, helping...  ...storage, and AI software into repeatable...  ...Product Architects, Engineering, Solution...  ...infrastructure from rack-scale systems through full...  ...SummaryLocation: US, CA, Santa Clara; US, MA, Westford... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $160k - $253k

     ...computing is the engine of artificial...  ...and a full-stack software ecosystem to power AI at scale. We are looking...  ...platforms, and rack-scale...  ...and rack-scale systems. This role bridges...  ...Explain critical infrastructure technologies like...  ...SummaryLocation: US, CA, Santa Clara; US, CA,... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $208k - $260k

     ...their hybrid cloud infrastructure. Gigamon has...  ...organizations. As a Principal Security Detections Engineer on the GigaSMART team...  ...high-performance systems software to identify suspicious...  ...based out of our Santa Clara, CA headquarters,...  ...strong focus on scale, accuracy, and performance... 
    Part time
    Local area
    Worldwide
    3 days per week

    Gigamon

    Santa Clara, CA
    2 hours ago
  • $208k - $260k

     ...their hybrid cloud infrastructure. Gigamon has served...  ...organizations. As a Principal Software Engineer, you will lead the design...  ...across distributed systems, backend services,...  ...is based out of our Santa Clara, CA headquarters, following...  ...mission to grow and scale an innovative... 
    Part time
    Local area
    Worldwide
    3 days per week

    Gigamon

    Santa Clara, CA
    2 hours ago
  • $272k - $431.25k

     ...to build planet-scale maps supporting self...  ...and large-scale systems. You will...  ...diverse team of engineers in mapping, perception...  ...simulation, planning, and infrastructure teams to...  ...production-quality software systems.Solid foundation...  ...: US, CA, Santa Clara; US, RemoteType:... 
    Full time
    Part time
    Worldwide

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $184k - $287.5k

     ...are seeking a Senior Software Engineer focused on container and cloud infrastructure. You will help design...  ...reliability, performance, and scale across thousands of...  ...Helm chart design systems, Operators, and platform...  ...: US, CA, Santa Clara; US, CA, RemoteType: Full... 
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    2 hours ago
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System Software team...  ...and managing workloads at production scale. Drive triage of the most difficult...  ...characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time... 
    Full time
    Part time
    Shift work

    Nvidia

    Santa Clara, CA
    7 hours ago
  • $210.16k - $271.98k

     ...Senior Principal Systems Development EngineerHelp...  ...deliver Dell's L11 rack-scale AI solutions:...  ...Systems Development Engineer, you'll own the...  ..., thermal, software, diagnostics, and...  ...most demanding AI infrastructure gets built and deployed...  ...team in Santa Clara, California.Essential... 
    Part time

    Dell Technologies

    Santa Clara, CA
    2 hours ago
  • $184k - $287.5k

     ...the clock speed of software development. We...  ...to build AI speed infrastructure for Tegra: a...  ..., and validation system for the future, aimed...  ...fundamentally a performance engineering role. You will be...  ..., and how we scale these next-...  ...: US, CA, Santa Clara; US, WA, RedmondType... 
    Full time
    Part time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    2 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure (Santa Clara). Be the first to apply!