Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Engineer - Rack Scale Systems Infrastructure

$272k - $431.25k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software systems. These systems support our upcoming rack-scale infrastructure products and services. This exceptional role sits where software meets hardware. You will work on control planes, state machines, orchestration systems, firmware, OS lifecycle, and networking fabrics. Your task is to compose infrastructure-as-a-service control plane software that converts complex rack-scale hardware into dependable, manageable, and programmable infrastructure for NVIDIA, partners, and leading cloud and enterprise clients globally.What You Will Be Doing:Define the complete software architecture for rack-scale infrastructure products and services, covering control plane services, infrastructure management, firmware, operating systems, kernel drivers, networking fabrics, accelerator software, and user-mode manageability software.Use Kubernetes and cloud-native primitives as an infrastructure fabric when appropriate. This includes controllers, operators, reconciliation loops, and open source components. These components can operate safely at rack and fleet scale. Build open source infrastructure software that can be embraced in different forms, including libraries, services, controllers, operators, and integration APIs for internal deployments and CSP environments.Bridge hardware and software teams across firmware, BMC, BIOS, boot flows, OS images, drivers, networking, NVLink domains, InfiniBand, GPUs, DPUs, CPUs, and system management interfaces. Translate forward-looking infrastructure roadmaps into formal software requirements, architecture specifications, and execution plans that align teams across the organization.Partner directly with hyperscalers, CSPs, enterprise customers, internal component leads, vendors, and business partners to align infrastructure capabilities with real-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies that reduce risk before hardware reaches production environments.Mentor senior engineers and technical leads, raising the engineering bar for large-scale networked systems, foundational software, and rack-scale control plane development.Make high-quality technical decisions in ambiguous environments, balancing customer needs, schedule, hardware realities, software maintainability, open source adoption, and long-term infrastructure evolution.What We Need To See:BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience. Proven experience (15+ years) in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering.Solid architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrade and rollback workflows, and distributed systems tradeoffs.Practical coding skills in Go, C++, or Rust, encompassing the capability to write, review, and direct production-quality infrastructure software. Experience with Rust is highly valued.Experience with Kubernetes or similar orchestration systems, especially as a fabric for managing infrastructure, hardware resources, or large-scale infrastructure services. Experience with Linux-based infrastructure software, OS rollout and image management, kernel or driver interactions, firmware lifecycle, and hardware bring-up workflows.Strong understanding of data center networking technologies and protocols, such as Ethernet, InfiniBand, RDMA, and fabric-level manageability. Experience with complex accelerator-based systems, including GPUs, DPUs, FPGAs, custom silicon, or other high-performance computing systems.Expertise in in-band and out-of-band management architectures, including BMCs, Redfish, IPMI, and related system management protocols. Ability to work with security experts to define practical tradeoffs across secure boot, attestation, access control, update safety, serviceability, and ease of operation.Experience crafting software intended for open source release, including API stability, modularity, documentation, community usability, and clean separation between shared software and deployment-specific integrations.Experience using AI-assisted development tools responsibly as an engineering multiplier for coding, test generation, debugging, build iteration, and documentation.Established skill in specifying requirements, guiding architecture, and managing delivery across various engineering teams and organizations. Strong written and verbal communication skills, enabling clear explanation of complex hardware/software tradeoffs to engineering leaders, customers, partners, and executives.Ways To Stand Out from the crowd:Built software supporting multiple adoption models — internal services, CSP-integrated offerings, reusable libraries, and customer-extensible APIs. Strong Rust skills in systems, infrastructure, or hardware-adjacent software.Multiplied team impact through reference implementations, design reviews, shared libraries, architecture docs, dev workflows, and AI-assisted engineering. Hands-on with fleet-scale provisioning, updates, rollback, observability, health, and remediation.Led across the full data center product lifecycle: inception, pre- and post-silicon, manufacturing, deployment, and operations. Familiar with open source ecosystems, contribution models, and balancing community collaboration with product needs.Deep experience with rack- or cluster-scale systems spanning compute, networking, storage, accelerators, firmware, and infra management as one operational domain. Skilled at finding simple, durable abstractions in complex systems to align teams, customers, and long-term direction.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 14, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US, Remote; US, MA, RemoteType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Software Engineer - Rack Scale Systems Infrastructure in Texas vacancy
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale. In this role... 
    Suggested
    Full time
    Remote work
    Shift work

    Nvidia

    Austin, TX
    1 day ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as...  ...point for GPU firmware and GPU system software, working directly...  ...NVIDIA GPU firmware at fleet scale. You will drive work streams...  ...across multiple GPUs in a rack-scale systemUnderstanding of... 
    Suggested
    Full time
    Remote work

    Nvidia

    Austin, TX
    1 day ago
  • $147k - $211k

     ...code developed by other engineers and provide feedback...  ...feedback. Triage product or system issues and debug/track...  ...performance, large scale systems data analysis,...  ...and resolution, and software test engineering....  ...built by the Technical Infrastructure team to keep it running... 
    Suggested
    Worldwide

    Google

    Sunnyvale, TX
    2 days ago
  • $224k - $356.5k

    NVIDIA is hiring engineers to scale up the introduction of next generation...  ...into its EDA Infrastructure. We expect you to have...  ...introductions (NPIs), distributed systems, familiarity with software testing and deployment,...  ...to join the EDA Team. Principal Software Engineer... 
    Suggested
    Full time

    Nvidia

    Austin, TX
    3 days ago
  • $152k - $241.5k

    Sr Software Engineer - Distributed Systems Engineer, EDA InfrastructureNVIDIA is hiring engineers to build and scale the infrastructure that supports our Electronic Design Automation (EDA) workloads. We are looking for engineers with strong programming skills, a deep understanding... 
    Suggested
    Full time
    Remote work

    Nvidia

    Austin, TX
    1 day ago
  • $135.2k - $306.4k

    Oracle Cloud Infrastructure (OCI) Compute delivers bare metal...  ...and GPU -at global scale. We are hiring a Consulting...  ...deliver distributed systems that are multi-tenant,...  ....Mentor and guide engineers in distributed systems...  ...IC5As a member of the software engineering division,... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    3 days ago
  •  ...digital, energy transition and infrastructure challenges. Job...  ...DetailsViridien is seeking a Platform Engineer - Infrastructure & Cloud Systems to design, build, and...  ...systems that support software deployment and execution across large-scale, distributed environments.... 
    Full time
    Relocation
    Flexible hours

    CGG

    Houston, TX
    4 days ago
  •  ...performing team that delivers infrastructure and performance...  ...a Lead Infrastructure Engineer at JPMorganChase...  ...apply deep knowledge of software, applications, and technical...  ...downstream data and systems or technical...  ...including cluster setup, scaling, and CI/CD pipeline integrationLeads... 
    Shift work

    JP Morgan Chase

    Plano, TX
    17 hours ago
  • $174k - $252k

    Write and test product or system development code....  ...code developed by other engineers and provide feedback...  ...maintaining, or launching software products, and 1 year...  ...at massive scale, and extend well beyond...  ...solutions.The AI and Infrastructure team is redefining what... 
    Worldwide

    Google

    Sunnyvale, TX
    2 days ago
  • $147k - $211k

     ...code developed by other engineers and provide feedback...  ...feedback. Triage product or system issues and debug/track...  ...and resolution, and software test engineering.In...  ...solutions.The AI and Infrastructure team is redefining what...  ...at unparalleled scale, efficiency, reliability... 
    Worldwide

    Google

    Sunnyvale, TX
    7 days ago
  • $61k - $101k

     ...mechanical or electrical engineering. We need strong...  ...understanding of production rack environments across...  ...AI tools to support infrastructure engineering workflows,...  ...them into end-to-end system designs. We drive continuous...  ...across hardware and software stakeholders. We... 
    Full time

    J.P. Morgan

    Houston, TX
    8 days ago
  • $162.6k - $244k

     ...Technologies, Inc. Job Area Engineering Group, Engineering...  ...Qualcomm Data Center AI System Hardware and Validation...  ...team develops rack-level AI inference solutions...  ...generation Qualcomm rack‑scale designs within a matrixed...  .... Drive firmware/software/hardware co‑design execution... 
    Work experience placement
    Work at office
    Work from home

    Qualcomm

    Austin, TX
    1 day ago
  • $272k - $431.25k

    We're looking for a Principal Engineer to join our CSP Engagements...  ...patterns and drive systemic improvements in...  ...the latest NVIDIA rack-scale systems, GPU architectures...  ...configuration, software, or workload differences...  ...work (profiling infrastructure, benchmark harnesses... 
    Full time
    Remote work

    Nvidia

    Austin, TX
    1 day ago
  • $145k - $165k

     ...Title: Senior. Systems Engineer - Infrastructure & Cloud Location: Houston, TX Salary: $145,000 - $165,000 / year Sponsorship: Position...  ...Experience designing, implementing, and supporting enterprise-scale Linux server platforms. ~ Extensive experience... 
    Full time
    Local area

    Addison Group

    Houston, TX
    25 days ago
  • $114.6k - $234.6k

    OCI (Oracle Cloud Infrastructure) AI Infrastructure is at the forefront...  ...AI revolution, creating systems that allow customers to scale from tens to thousands...  ...Looking For:Adaptable Engineers: Self-motivated...  ...of the stack, as well as software debugging and low-level... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    3 days ago
  •  ...performing team that delivers infrastructure and performance...  ...a Lead Infrastructure Engineer at JPMorganChase...  ...apply deep knowledge of software, applications, and technical...  ...downstream data and systems or technical...  ...including cluster setup, scaling, and CI/CD pipeline integrationLeads... 
    Shift work

    JP Morgan Chase

    Plano, TX
    8 days ago
  • $184k - $287.5k

     ...the world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely...  ...performance practices in large-scale GPU infrastructure, delivering powerful tools, methodologies... 
    Full time
    Remote work

    Nvidia

    Texas
    4 days ago
  • $107.5k - $204.5k

     ...experience and renowned engineering expertise to meet the...  ...-domain C2 and radar systems. The center specializes...  ..., and Radar Interface software for Tactical Radar...  ...an opportunity for a Principal Software Engineer to join...  ...architecture for large scale and complex system‑of‑... 
    Temporary work
    Work experience placement
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    Raytheon

    Plano, TX
    1 day ago
  • $105.1k - $189.2k

     ...offering comprehensive engineering, supply chain, and...  ...Title: Lead Photonics System EngineerJabil is seeking...  ...elements at the box and rack level at the EMS...  ...within the Intelligent Infrastructure business unit. Lead photonics...  .... This is production scale photonics... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide

    Jabil Circuit

    Austin, TX
    1 day ago
  • $184k - $287.5k

    NVIDIA is looking for a Systems & Software Engineer interested in building and running reliable large scale infrastructure platform services. In this role you will create and support automation that manages infrastructure, as part of our organization running EDA platforms... 
    Full time
    Remote work

    Nvidia

    Austin, TX
    17 hours ago
  • $114.6k - $234.6k

     ...strong understanding of operating systems, computer networks, and high-...  ...and operating large-scale production systems with 1,000...  ...and experience across the full software development lifecycle. Preferred...  ...: We are Oracle OCI AI Infrastructure, and we are building a... 
    Full time

    Oracle

    Austin, TX
    18 days ago
  • $159k - $231k

    Design system-level power delivery network...  ...degree in Electrical Engineering, Computer...  ...within Platforms Infrastructure Engineering, you will...  ..., package, board, rack, and pod levels. Your...  ...Infrastructure at unparalleled scale, efficiency,...  ...the future. From software to hardware our... 
    Worldwide

    Google

    Sunnyvale, TX
    2 days ago
  • $165k - $200k

     ...world-class team of scientists, engineers, and business professionals...  .... We are seeking a Principal Software Engineer to help design and...  ...the high-performance control systems that drive our next-generation...  ...quantum processors operate, scale, and perform. This is a hands... 
    Temporary work
    Work at office
    3 days per week

    Atom Computing

    Austin, TX
    27 days ago
  • $230k - $288k

     ...adding hardware, installing software, or changing a line of code....  ...the Department Cloudflare’s engineers build and operate the software...  ...build high‑growth products, help scale our expanding network, build...  ...and response times, and make systems failure‑resistant and ready‑... 
    Temporary work
    Flexible hours

    Cloudflare Inc

    Austin, TX
    2 days ago
  • $174k - $252k

    Write and test product or system development code....  ...solutions, leverage ML infrastructure, and evaluate tradeoffs...  ..., or launching software products, and 1 year of...  ...technologies.Google's software engineers develop the next-...  ...information at massive scale, and extend well beyond... 
    Worldwide

    Google

    Sunnyvale, TX
    4 days ago
  • $165.2k - $223.6k

     ...Labs designs silicon and software that accelerates...  ...clocks and voltages, to systems level device drivers for I2C infrastructure pervasive in the server...  ...help the organization scale though the use of software...  ...team members develop your engineering expertise so you feel... 
    Internship
    Local area
    Flexible hours

    Amazon

    Austin, TX
    17 hours ago
  •  ...enterprises with Infrastructure as a Service (IaaS...  ...(PaaS), and Software as a Service (SaaS...  ...automation and tooling engineering team within...  ...build the software systems that allow network...  ...infrastructure at scale, across dozens of...  ...from initial rack bring-up to day-to... 
    Full time
    Contract work
    Part time
    Fixed term contract
    Internship
    Shift work

    IBM

    Dallas, TX
    2 days ago
  • $159k - $230k

    Gather system requirements, define architecture, execute hardware...  ...functionally with Hardware, Software, Mechanical, Thermal,...  ...Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics...  ...data center.Our Platforms Infrastructure Engineering team designs and... 
    Worldwide

    Google

    Sunnyvale, TX
    4 days ago
  •  ...Principal Systems/Software Engineer Cloud & On-Premise Software Testing This role has been designed as 'Hybrid' with a requirement that you...  ...transitioning to a secure, cloud-enabled, mobile-friendly infrastructure. Many rely on a combination of both. Wherever they are... 
    Work at office
    Local area
    Relocation package
    2 days per week

    Jobleads-US

    San Juan, TX
    1 day ago
  •  ...come to the right place.As a Principal Software Engineer-KYC Risk Assessment at...  ...high-impact software that scales.Job responsibilitiesArchitects...  ...and governs agentic AI systems, including multi-agent workflows...  ...and model serving infrastructure or managed endpoints (AWS... 

    JP Morgan Chase

    Houston, TX
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Engineer - Rack Scale Systems Infrastructure. Be the first to apply!