Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Network Reliability Engineer - DGX Cloud

$136k - $224.25k

NVIDIA

NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence.

In this role, the Senior Network Operations Engineer will remediate critical alerts within defined SLAs, triage production impacting network incidents, and interact with internal customers on network related issues. They will also be responsible for engaging with external vendors to remediate hardware and software issues, and participate in project related work such as network device upgrades and capacity augmentations. An ideal candidate will possess a wide range of skills, including alert monitoring & resolution in large-scale networks and CSP environments, outstanding troubleshooting skills, understanding of L3 underlay networks, and network protocol knowledge in large multi-vendor infrastructures.

What you will be doing:

  • Engage in 24/7 global shift rotations to provide remote support for network repairs and changes while collaborating across teams and updating customers on status and ticket information.

  • Drive operational improvements in change management and daily operations by following procedures.

  • Manage and operate large scale IP network technologies and infrastructures.

  • Utilize your skills in Peering and Datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits.

  • Monitor and support the network health of on-premises and cloud infrastructures.

  • Collaborate and develop workflow enhancements while documenting best practices.

What we need to see:

  • Deep knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec.

  • 5+ years of experience in network operations.

  • Skilled in network troubleshooting techniques and demonstrating creative problem-solving abilities.

  • Strong track record of alert response within defined SLAs and Incident management.

  • Experience with one or more of the following CSP environments: AWS, Azure, GCP, OCI.

  • Familiarity with Arista, Fortinet and Juniper.

  • Hands-on experience with contributing to tooling and automation for provisioning, monitoring, and managing complex network infrastructures.

  • Bachelor’s degree in Computer Science, related technical field, or equivalent experience.

  • Excellent verbal and written communication skills.

Ways To Stand Out From The Crowd:

  • Solid understanding of Mellanox/Cumulus OS and Infiniband technology.

  • Skilled in Unix/Linux system administration, with the ability to write and understand Python/Shell scripts to improve efficiency in hyperscale environments.

  • Familiarity with leveraging tools such as Netbox/Nautobot, Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate about innovating and investing in ground breaking technologies.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these outstanding opportunities in deep learning cloud solutions.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 136,000 USD - 224,250 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4.

You will also be eligible for equity and benefits ( .

Applications for this job will be accepted at least until May 22, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Network Reliability Engineer - DGX Cloud in Santa Clara, CA vacancy
  • $136k - $224.25k

    ## Senior Network Reliability Engineer - DGX CloudApplylocations: US, CA, Santa Clara: US, Remotetime type: Full timeposted on: Posted Yesterdayjob requisition...  ...Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network... 
    Senior
    Network
    Remote work
    Shift work

    NVIDIA Corporation

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure...  .... We are looking for Senior Software Engineers to help build the...  ...systems that make GPU clusters reliable, scalable, and safe to run...  ...with platform, storage, networking, security, and workload teams... 
    Senior
    Network
    Remote work

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...the world. As part of the DGX Cloud organization, the...  ...security, silicon, and cloud engineering teams to turn embedded hardware...  ...attestation standards into reliable, self-service cloud capabilities...  ...across Data Center, Automotive, Networking, and AI ecosystems.... 
    Senior
    Network
    Remote work

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $320k

    Director, Site Reliability and Software Engineering - DGX Cloud page is loaded## Director, Site Reliability and Software Engineering - DGX Cloudlocations:...  ...researchers can now rapidly build, train, and deploy neural network models to address some of the most complicated AI... 
    Network

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

    NVIDIA Corporation is seeking a Senior Network Engineer to develop a cloud network infrastructure that supports software development workflows. This role involves designing, implementing, and troubleshooting network stacks, with a focus on automation. Key qualifications... 
    Senior
    Network

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

    NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient network to support NVIDIA software development workflows and tools, including CI/CD pipelines, compute resource management... 
    Senior
    Network

    NVIDIA Corporation

    Santa Clara, CA
    1 day ago
  • $136k - $264.5k

    NVIDIA Corporation is seeking a Senior Network Reliability Engineer to support and maintain cloud and datacenter network infrastructures in Santa Clara, California. The role entails engaging in 24/7 global support, managing large scale IP network technologies, and working... 
    Senior
    Network

    NVIDIA Corporation

    Santa Clara, CA
    1 day ago
  • $176k - $276k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain...  ...knowledge across different systems, networking, coding, database, capacity management...  ...and deployment and open source cloud enabling technologies like Kubernetes... 
    Senior
    Network

    NVIDIA Corporation

    Santa Clara, CA
    5 days ago
  • $272k - $431.25k

     ...NVIDIA DGX Cloud is scaling GPU infrastructure across...  ...for Principal Software Engineers to help shape the technical...  ..., automation, and reliability across large-scale GPU...  .... This role is for senior technical leaders who...  ...infrastructure, storage, networking, security, and... 
    Network

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $80 per hour

     ...Senior Cloud DevOps Engineer/Site Reliability Engineer Position Title: Senior Cloud DevOps Engineer/Site Reliability Engineer Location: San Jose,...  ..., CI/CD, automated testing) Good understanding of networking Bachelor degree in Computer Science or equivalent... 
    Senior
    Network
    Local area

    ClifyX

    San Jose, CA
    1 day ago
  • $90k - $215k

     ...Senior Software Engineer- Observability and Reliability Platform Engineering (REMOTE) Senior Software Engineer- Observability and Reliability Platform Engineering...  ..., and maintenance of the hardware, software, and network systems 3+ years of experience in open-source... 
    Senior
    Network
    Hourly pay
    Full time
    Work experience placement
    Local area
    Remote work
    Flexible hours

    GEICO

    San Jose, CA
    5 days ago
  • NVIDIA Corporation is seeking a Senior Software Engineer to join its DGX Cloud Production Engineering team in Santa Clara, CA. This role focuses on building...  ...systems for large-scale GPU clusters, ensuring reliability and scalability. The ideal candidate will have over... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 
    Senior
    Network

    TechDigital Group

    Santa Clara, CA
    5 days ago
  • $120k - $243k

    Hewlett Packard Enterprise Development LP is seeking a Senior Competitive Technical Marketing Engineer in Sunnyvale, California. This is an onsite role focused on providing competitive analysis of the HPE Networking portfolio. The ideal candidate will assist in the... 
    Senior
    Network

    Hewlett Packard Enterprise Development LP

    Sunnyvale, CA
    5 days ago
  • $147k - $237.5k

    Palo Alto Networks, Inc. is seeking a Senior Performance Engineer based in Santa Clara, California. The role includes designing and implementing performance testing strategies for distributed cloud-native systems and driving resolution of performance-related issues across... 
    Senior
    Network

    Palo Alto Networks, Inc.

    Santa Clara, CA
    2 days ago
  • $134k - $215.5k

    Palo Alto Networks, Inc. is seeking a Senior Technical Marketing Engineer to join their Cloud-Delivered Security Services team in Santa Clara, California. This role involves developing technical marketing strategies, creating authoritative content, and delivering high-impact... 
    Senior
    Network

    Palo Alto Networks, Inc.

    Santa Clara, CA
    5 days ago
  • Palo Alto Networks, Inc. is seeking a Senior Staff Engineer to contribute to their innovative cloud security product, Data Loss Prevention (DLP). This role involves utilizing backend Java cloud engineering skills to develop a cutting-edge, industry-leading service aimed... 
    Senior
    Network
    Work at office
    3 days per week

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 days ago
  • $264.51k - $298.62k

    Versa Networks in Santa Clara focuses on developing cutting-edge software solutions for Network Management Systems. They are seeking a Software Engineer with expertise in Java and cloud technologies. The ideal candidate will work on the development of modules for the Versa... 
    Senior
    Network
    Remote job

    Versa Networks

    Santa Clara, CA
    2 days ago
  • $140k - $215k

     ..., Inc. is looking for a backend software engineer to join their Ingestion group in Sunnyvale...  ...position involves managing high-volume network communications and requires over 7 years...  ...with experience in distributed systems and cloud technologies. The role offers high autonomy... 
    Senior
    Network

    CrowdStrike Holdings, Inc.

    Sunnyvale, CA
    1 day ago
  • $120.3k - $194.53k

    Our Mission At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the...  ...a large hybrid infrastructure across multiple public clouds. As a Site Reliability Engineer on the Internet Security Platform team, you will be part... 
    Senior
    Network
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks, Inc.

    Santa Clara, CA
    5 days ago
  • $104k - $169.5k

    Palo Alto Networks, Inc. is looking for an experienced Backend Engineer to design, develop, and deliver next-generation technologies. You will work on scalable software...  .... This role requires strong programming skills, cloud technology expertise, and the ability to thrive in... 
    Senior
    Network

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the...  ...record debugging large-scale, cloud-native stacks across networking (RDMA/RoCE), storage, and control planes. ~... 
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    12 hours ago
  • $184k - $356.5k

    NVIDIA Corporation is looking for a Senior Systems Software Engineer in Santa Clara, CA. This role involves designing and enhancing networking management services, building UI components, and collaborating with various teams to support software-defined and data center... 
    Senior
    Network

    NVIDIA Corporation

    Santa Clara, CA
    5 days ago
  •  ...leading cybersecurity firm is seeking a Senior Backend Software Engineer to focus on the Azure Firewall...  ...in Go / Golang and familiarity with cloud environments like AWS or Azure. You...  ...into existing products and focus on networking aspects. The role is hybrid, requiring... 
    Network
    Work at office

    Illumio

    Sunnyvale, CA
    5 days ago
  •  ...cybersecurity company in Sunnyvale is seeking a Senior Backend Software Engineer with strong Go/Golang coding skills and cloud experience in Azure or AWS. In this hybrid...  ...firewall management security frameworks and work on networking aspects of innovative products. Join a... 
    Senior
    Network

    Illumio

    Sunnyvale, CA
    1 day ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...customer-facing reliability.* Strong Linux + networking fundamentals, distributed systems... 
    Senior
    Network

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  •  ...leading cybersecurity company is seeking a Senior Backend Software Engineer to join their team in Sunnyvale, CA. In...  ...ecosystem using Go / Golang. Experience with cloud environments, particularly Azure or AWS, as well as networking and API development, is essential. Join a... 
    Senior
    Network

    Illumio

    Sunnyvale, CA
    2 days ago
  • $224k - $356.5k

     ...service powered by NVIDIA GPUs in the cloud, transforms a Mac, any PC, or just a...  ...Visit us at We are looking for a Senior Systems Software engineer to join a team of highly skilled and...  ...algorithms and solve complicated networking and routing issues? Are you interested... 
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $174k - $252k

    Senior Software Engineer, Site Reliability Engineering X Applicants in San Francisco: Qualified applications with arrest or conviction records will be...  ...taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and... 
    Senior
    Network
    Full time

    Google Inc.

    Sunnyvale, CA
    4 days ago
  • $176k - $276k

     ...intelligence. Join our team of innovative engineers who develop and maintain software...  ...understanding of operating systems, computer networks, and high-performance applications. ~ Proven...  ...-focused hardware and software, such as DGX systems and Compute Clusters.... 
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Network Reliability Engineer - DGX Cloud. Be the first to apply!