Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Network Reliability Engineer - DGX Cloud

$136k - $264.5k
Full-time

NVIDIA

NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence.

In this role, the Senior Network Operations Engineer will remediate critical alerts within defined SLAs, triage production impacting network incidents, and interact with internal customers on network related issues. They will also be responsible for engaging with external vendors to remediate hardware and software issues, and participate in project related work such as network device upgrades and capacity augmentations. An ideal candidate will possess a wide range of skills, including alert monitoring & resolution in large-scale networks and CSP environments, outstanding troubleshooting skills, understanding of L3 underlay networks, and network protocol knowledge in large multi-vendor infrastructures.

What you will be doing:

  • Engage in 24/7 global shift rotations to provide remote support for network repairs and changes while collaborating across teams and updating customers on status and ticket information.

  • Drive operational improvements in change management and daily operations by following procedures.

  • Manage and operate large scale IP network technologies and infrastructures.

  • Utilize your skills in Peering and Datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits.

  • Monitor and support the network health of on-premises and cloud infrastructures.

  • Collaborate and develop workflow enhancements while documenting best practices.

What we need to see:

  • Deep knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec.

  • 5+ years of experience in network operations.

  • Skilled in network troubleshooting techniques and demonstrating creative problem-solving abilities.

  • Strong track record of alert response within defined SLAs and Incident management.

  • Experience with one or more of the following CSP environments: AWS, Azure, GCP, OCI.

  • Familiarity with Arista, Fortinet and Juniper.

  • Hands-on experience with contributing to tooling and automation for provisioning, monitoring, and managing complex network infrastructures.

  • Bachelor’s degree in Computer Science, related technical field, or equivalent experience.

  • Excellent verbal and written communication skills.

Ways To Stand Out From The Crowd:

  • Solid understanding of Mellanox/Cumulus OS and Infiniband technology.

  • Skilled in Unix/Linux system administration, with the ability to write and understand Python/Shell scripts to improve efficiency in hyperscale environments.

  • Familiarity with leveraging tools such as Netbox/Nautobot, Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate about innovating and investing in ground breaking technologies.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these outstanding opportunities in deep learning cloud solutions.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 136,000 USD - 224,250 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 23, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Network Reliability Engineer - DGX Cloud in Santa Clara, CA vacancy
  • $168k - $270.25k

     ...excellent, comprehensive support to our customers! ​Sr Site Reliability Engineer in this role will significantly impact and contribute to...  ...Manager clusters is a definite plus.Proficiency with cluster networking including InfiniBand and Spectrum-XNVIDIA is widely considered... 
    Senior
    Network
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing...  ...end behavior across GPUs, networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,...  ...raise the performance and reliability of AI workloads. Join our... 
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $200k - $322k

     ...NVIDIA’s DGX Cloud team helps some of the most advanced AI builders in the world move from...  ...recommendations across compute, networking, storage, and cloud environments.Turn repeat...  ...and reduce friction.gainsightWork across Engineering, Product, Operations, and Finance to... 
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure...  .... We are looking for Senior Software Engineers to help build the...  ...systems that make GPU clusters reliable, scalable, and safe to run...  ...with platform, storage, networking, security, and workload teams... 
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $287.5k

     ...As part of DGX Cloud, the Attestation & Trust Services...  ...scale. We want hands-on engineers with strong systems...  ...production support. Improve reliability and security in the...  ..., working with senior engineers on problems...  ...practical knowledge of Linux, networking, APIs, concurrency,... 
    Senior
    Network
    Full time
    Local area
    Remote work

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $184k - $356.5k

     ...Joining NVIDIA's DGX Cloud Lepton Team means contributing...  ...software engineer to join our team. You'...  ...in production. As a senior DGX Cloud AI Infrastructure...  ...meaningful and actionable reliability metrics to track and improve...  ...of NVIDIA GPUs, network technologies (RDMA, IB... 
    Senior
    Network
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  ...data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability,...  ...corrective actions.Partner with Compute, Networking, Storage, Security, and Support teams... 
    Senior
    Network
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $200k - $322k

    ## Senior Customer Success Engineer - DGX CloudApplylocations: US, CA, Santa Clara: US, Remotetime type: Full...  ...Todayjob requisition id: JR2022184The DGX Cloud organization bridges customer...  ...scale — spanning compute, storage, networking, and GPU capacity management across... 
    Senior
    Network

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

    NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient network to support NVIDIA software development workflows and tools, including CI/CD pipelines, compute resource management... 
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $224k - $356.5k

     ...is transforming how the world uses AI, cloud, and accelerated computing, and trust is...  ...and cloud teams to bring new ideas into reliable production services that people rely on...  ...NVIDIA platforms across data center, AI, networking, and partner environments.Improving reliability... 
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage,...  ...art LLM workloads run efficiently and reliably at scale. You will lead deep performance...  ...performance across compute, memory, networking, and communication layers using tools... 
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $166k - $244k

    Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run...  ...systems. SRE ensures that Google Cloud's services—both our internally critical...  ...apart so we can rebuild them. We keep our networks up and running, ensuring our users... 
    Senior
    Network
    Full time

    Google

    Sunnyvale, CA
    2 days ago
  • $152k - $287.5k

     ...NVIDIA is seeking a Senior Software engineer to build the next generation of our...  ...operate, and scale across cloud and on-premises environments...  ...Design and build highly reliable distributed systems and APIs...  ...spanning infrastructure, runtime, networking, hardware, and operations,... 
    Senior
    Network
    Full time

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based...  ...across environments.We are looking for a hands-on senior engineer to own the lifecycle and automation of the... 
    Senior
    Network
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $168k - $264.5k

     ...transform the way people work and play. NVIDIA is seeking a Senior Network Deployment Engineer to help build and scale our global network. In this role,...  .../training a strong plusExperience with multi-cloud networking (AWS VPC, Azure VNet, GCP VPC)Excellent architecture... 
    Senior
    Network
    Full time
    Contract work
    Work experience placement
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $176k - $276k

    Production engineering is a field that involves crafting, building, and...  ..., data management, systems, networking, coding, database management,...  ...deployment, along with open-source cloud-enabling technologies such as...  ...storage architectures are reliable, scalable, and efficient.... 
    Senior
    Network
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    2 days ago
  • $90k - $215k

     ...Senior Software Engineer- Observability and Reliability Platform Engineering (REMOTE) Senior Software Engineer- Observability and Reliability Platform Engineering...  ..., and maintenance of the hardware, software, and network systems ~3+ years of experience in open-source... 
    Senior
    Network
    Hourly pay
    Full time
    Work experience placement
    Local area
    Remote work
    Flexible hours

    GEICO

    San Jose, CA
    16 hours ago
  • $272k - $431.25k

     ...NVIDIA DGX Cloud is scaling GPU infrastructure across...  ...for Principal Software Engineers to help shape the technical...  ..., automation, and reliability across large-scale GPU...  ....This role is for senior technical leaders who...  ...infrastructure, storage, networking, security, and workload... 
    Network
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...can make a lasting impact on the world.Are you passionate about building world-class reliability systems? Join NVIDIA as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a pivotal part of a team that redefines operational excellence. Our team is at... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...the DSX Kubernetes Fleet team within NVIDIA's DGX Cloud organization - a collaborative group of cloud platform engineers, architects, and SREs who are passionate...  ....Understanding of performance, security, and reliability in complex distributed systems.Ways to stand... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $264.5k

     ...outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at...  ...resolution across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the...  ..., increase efficiency in our use of cloud resources and our developer’s time,...  ...stackDeep understanding of AWS Networking, Compute, Storage, and managed services... 
    Senior
    Network
    Flexible hours

    Sumo Logic

    San Jose, CA
    11 hours ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...Lambda’s multi-tenant cloud networking platform and SDN infrastructureOperate...  ...teams to improve service reliability and deployment... 
    Senior
    Network
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $210.6k - $305.1k

     ...deliver flawless digital experiences across every network - even the ones they don’t own. Powered by AI and an unmatched set of cloud, internet and enterprise network telemetry...  ...:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for... 
    Senior
    Network
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    4 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...so we can rebuild them. We keep our networks up and running, ensuring our users have... 
    Senior
    Network

    Google

    Sunnyvale, CA
    2 days ago
  •  ..., and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP...  ...safely and predictably.Cloud Security: Harden our network architecture and application security posture,... 
    Senior
    Network
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $90k - $180k

     ...than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our...  ..., including distributed systems and cloud deployments in Azure. Work closely with...  ...design patterns.Solid Linux & networking fundamentals — DNS, TCP/IP, TLS, load... 
    Senior
    Network
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  • $267k - $356k

    Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...the industry, which means reliability and performance aren't just...  ...etc.Work with hardware and networking teams to diagnose low-level... 
    Senior
    Network
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $101k - $161k

    Company DescriptionArista Networks is an industry leader...  ...-driven, client-to-cloud networking for large data...  ...awards, such as Best Engineering Team, Best Company for...  ...’re looking for Site Reliability Engineers to join our...  ...EngineeringExperience level: Mid-Senior LevelIndustry:... 
    Senior
    Network

    Arista Networks

    Santa Clara, CA
    11 hours ago
  •  ...Mirantis empowers platform engineering teams to deliver...  ...environment—on-premises, in the cloud, at the edge, or in...  ...to GPU architecture, networking, and orchestration....  ...reference-system families (DGX, HGX, MGX); aware of what...  .... Indicative OTE: senior people-leader band for... 
    Senior
    Network
    Full time
    Remote work

    Mirantis

    San Jose, CA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Network Reliability Engineer - DGX Cloud. Be the first to apply!