Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer - NVLink Rack Scale Stability and Reliability

$152k - $241.5k

NVIDIA

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.We are looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus on NVLink Rack-Scale Systems Stability & Reliability. In this role, you will partner closely with architects and developers building our next-generation NVLink and NVSwitch systems, helping transform first-of-their-kind platforms into stable, reliable, and volume production-ready systems. You will work on complex system-level challenges spanning resiliency, diagnostics, recovery, and large-scale AI infrastructure, contributing directly to the software foundation powering next-generation datacenter deployments. What you will be doing:Drive platform bringup, feature enablement, end-to-end software validation, and debug for next-generation NVLink-based GPU and rack-scale systems.Develop tools, diagnostics, automation, and infrastructure for system validation, regression testing, and fleet support.Lead reliability and MTBI validation through stress testing, telemetry analysis, failure injection, and issue resolution.Triage complex software, firmware, networking, and platform issues across validation, deployment, and production environments.Collaborate with architecture, hardware, firmware, software, and Customer engagement teams to improve system quality and reliability.Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and operational readiness.Create automation, dashboards, runbooks, and debug workflows that improve root-cause analysis and operational efficiency.What we need to see:BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.Strong programming skills in C/C++ and Python; Bash/Shell scripting experience is a plus.Strong system-level debugging across software, firmware, hardware, and networking layers.Solid networking fundamentals, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.Experience with large-scale AI systems, including platform bringup, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging.Ability to triage complex multi-domain issues using logs, telemetry, experiments, and structured debugging methods.Strong communication and collaboration skills across engineering, customer, and operations teams.Passion for building reliable next-generation AI infrastructure and solving complex system-level challenges at scale.Ways to stand out from the crowd:Experience with NVIDIA GPU systems, NVLink, NVSwitch, CUDA, and large-scale AI/HPC clusters such as NVIDIA GB200 NVL72.Strong understanding of large-scale AI system architecture, including PCIe, memory hierarchy, DMA, high-speed interconnects, and distributed training/inference systems.Experience with server management technologies, data center operations, cluster provisioning, scaling, and fleet monitoring.Proven experience building diagnostics, automation, CI/CD pipelines, dashboards, and reliability tooling.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 31, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, IL, Remote; US, CO, Remote; US, AZ, Remote; US, CA, Remote; US, MA, RemoteType: Full time

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer - NVLink Rack Scale Stability and Reliability in Santa Clara, CA vacancy
  • $272k - $431.25k

     ...as a Principal Rack Scale Systems Infrastructure Engineer, you will build...  ...development of software systems. These...  ...drivers, networking, NVLink domains,...  ...needs. Establish reliability, security, validation...  ....Mentor senior engineers and technical...  ...including API stability, modularity,... 
    Suggested
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...'re looking for a Principal Software Engineer to join our CSP Engagements...  ...the technical focal point for rack-scale system SW/FW, working with...  ..., and operate these systems reliably at fleet scale. In this role...  ...power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking... 
    Suggested
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...re looking for a Principal Software Engineer to join our CSP Engagements...  ...technical focal point for fleet-scale reliability, working directly with...  ...Deep expertise in multi-NUMA, rack-scale system software and...  ...taxonomy (Xid errors, NVLink error counters, thermal events... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...to large multi-node NVLink domain rack architectures. These...  ...optimized NVIDIA AI and HPC software stack. We’re...  ...operationalize rack-scale factory and deployment...  ...passion for building reliable, debuggable, and scalable...  ...Mentor architects and engineering teams to grow them... 
    Senior
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $174k - $253k

     ...design consulting, developing software platforms and frameworks,...  ...latency and overall system health.Scale systems sustainably through...  ...for changes that improve reliability and velocity.Practice...  ...degree in Computer Science, Engineering, a related field, or equivalent... 
    Senior

    Google

    Sunnyvale, CA
    7 hours ago
  • $101k - $161k

     ...intelligence, and software-defined...  ...awards, such as Best Engineering Team, Best Company...  ...looking for Site Reliability Engineers to join...  ...production systems at scale. We are...  ...reliability, and stability. You’ll have firsthand...  ...EngineeringExperience level: Mid-Senior LevelIndustry:... 
    Senior

    Arista Networks

    Santa Clara, CA
    7 hours ago
  • $320k

     ...to large multi-node NVLink domain rack architectures. These...  ...optimized NVIDIA AI and HPC software stack. We're...  ...leader to drive the engineering roadmap and innovation...  ...architecture for NVIDIA's rack-scale productsMaintain deep...  ...high quality & reliable software; serving as... 
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    NVIDIA NVLink team is seeking a Senior Software Developer or manager to serve as Tech Lead...  ...infrastructure at scale. What you will be doing:...  ...product, test, applications engineering, production/manufacturing...  ...building, code quality, and reliability. Proven track record of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...times a day - quickly, reliably, and securely. Any...  ...an impact on a global scale, come make a difference...  ...Fiserv.Job TitleSenior Software Engineer - Integrated Payment Experience...  ...POS terminals. As a Senior Software Engineer, you...  ...to-end quality, fleet stability, and release... 
    Senior
    Full time
    Work at office
    Worldwide
    Monday to Friday

    Fiserv

    Sunnyvale, CA
    7 hours ago
  • $152k - $241.5k

     ...is built. We are seeking a Senior Software Engineer - AI Inference to advance open...  ..., low‑latency inference at scale.This is a hands-on role for...  ...inference performance and reliability: parallelism strategies,...  ...bandwidth, kernel fusion, PCIe/NVLink effects, and network... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...the world.Join a team that analyzes large-scale datacenter workloads on GPU-accelerated...  ...with OS, container, GPU, and systems engineers. When useful, you will apply machine learning...  .../prediction) inside existing software workflows.What we need to see:5+ years analyzing... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $117k - $234k

     ...you'll do...We’re seeking a Software Engineer to build personalized, high-...  ...enhance system performance, stability, and efficiencyEnsure secure...  ...transactional systems with strict reliability, consistency, and latency...  ...NoSQL data stores in high-scale production environments.... 
    Senior
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    7 hours ago
  • $117k - $234k

     ...included in this role As a Senior Software Engineer - Android, you will design, build, and scale customer‑facing mobile...  ...accessibility, testing, and production reliability through strong engineering...  ..., root‑cause analysis, and stability improvements.Ensure performance... 
    Senior
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    2 days ago
  • NVIDIA Corporation is seeking a candidate to analyze large-scale datacenter workloads on GPU-accelerated clusters. Responsibilities include identifying application improvements and building visualizations for data analysis. The ideal candidate has 5+ years of experience... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $184k

     ...world. We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale. At NVIDIA, this role is uniquely...  ...Decompose high-complexity performance or stability issues into minimal reproduction cases,... 
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    4 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the...  ...operational excellence of Sumo’s planet-scale observability and security products....  ...modern approaches to cloud-native software securityExperienced with agile frameworks... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    7 hours ago
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes...  ...plane services and dataplane software running on SmartNICsDevelop...  ...networking teams to improve service reliability and deployment workflowsDeploy... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  •  ...automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud...  ...automated architectures that will support our next 10x of scale. You will be the primary authority on reliability, performance... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $262k - $365k

     ...design consulting, developing software platforms and frameworks,...  ...latency and overall system health.Scale systems sustainably through...  ...for changes that improve reliability and velocity.Practice...  ...degree in Computer Science or Engineering.Site Reliability Engineering... 
    Senior

    Google

    San Jose, CA
    7 hours ago
  • $90k - $180k

     ...than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our...  ...eliminate performance bottlenecks in software and infrastructure, ensuring low-latency...  ...insights into system health and behavior at scale. Automate away manual operational... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  • $148k - $235.75k

     ....Join our team of innovative engineers who are building an AI Data Center...  ..., high-volume telemetry into reliable, job-centric insights and...  ...depend on. You’ll partner Software Engineering and Systems Engineering...  ...(deploying, debugging, scaling) for telemetry-heavy microservices... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $159.2k - $301.6k

     ...Graphs on the cloud. In this reliability-focused role, you will own the...  ...'ll partner with the backend engineers building these APIs to make...  ...resilient, and cost-effective as it scales. What you'll do: Define and...  ..., infrastructure, or backend software development with a strong... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    3 days ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our Compute Farm...  ....Experience supporting large‑scale HPC clusters using Slurm, LSF...  ...lifecycle management, fleet reliability/auto-healing, E2E...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...revolution, building the software and systems that...  ...We are looking for a Senior Software Engineer to lead the bring-up,...  ...platforms at the largest scales we run.In this role...  ...run efficiently and reliably at scale. You will lead...  ...topology, including NVLink, NVSwitch, PCIe, RoCE... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $139k - $257.55k

     ...impact on the customer experience. As a Senior Software Engineer focused on Data Science and Platform...  ...-latency platform that serves them at scale across Adobe's most sensitive, highest...  ...new surfaces turnkey. Keep it fast, reliable, and scalable as usage grows. • Partner... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide
    Shift work

    Adobe

    San Jose, CA
    1 day ago
  • $262k - $365k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...or Engineering.Site Reliability Engineering (SRE) combines...  ...to build and run large-scale, massively distributed, fault...  ...Engineer chose to join SRE.As the Senior Engineering Manager for... 
    Senior

    Google

    Sunnyvale, CA
    4 days ago
  • $117k - $234k

     ...What you'll do... Role summary: As a Senior Software Engineer at Walmart, you will lead the...  ...team develops and operates enterprise-scale Operational Excellence platforms that...  ...AI-enabled reports, the team enhances reliability, reduces operational risk, and accelerates... 
    Senior
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    4 hours ago
  • $136k - $224.25k

    NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter...  ...network serves the needs across the whole software stack for NVIDIA, from Graphics...  ...monitoring & resolution in large-scale networks and CSP environments, outstanding... 
    Senior
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...bill. Role Description As a Senior iOS Software Engineer on the Customer Experience...  ...to ensure a responsive and reliable application. Develop and maintain...  ...test suites to ensure app stability and high code quality....  ...Experience working with large-scale consumer-facing mobile... 
    Senior
    Full time

    GFiber

    Sunnyvale, CA
    1 day ago
  • $174k - $253k

     ...years of experience with software development in one or...  ...with developing large-scale infrastructure,...  ...projects.Google's software engineers develop the next-generation...  ...and maintaining a reliable, cost-efficient...  ...depth, a focus on system stability, and a commitment to making... 
    Senior

    Google

    Sunnyvale, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer - NVLink Rack Scale Stability and Reliability. Be the first to apply!