Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Network Reliability Engineer - DGX Cloud

$136k - $264.5k
Full-time

NVIDIA

NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence.

In this role, the Senior Network Operations Engineer will remediate critical alerts within defined SLAs, triage production impacting network incidents, and interact with internal customers on network related issues. They will also be responsible for engaging with external vendors to remediate hardware and software issues, and participate in project related work such as network device upgrades and capacity augmentations. An ideal candidate will possess a wide range of skills, including alert monitoring & resolution in large-scale networks and CSP environments, outstanding troubleshooting skills, understanding of L3 underlay networks, and network protocol knowledge in large multi-vendor infrastructures.

What you will be doing:

  • Engage in 24/7 global shift rotations to provide remote support for network repairs and changes while collaborating across teams and updating customers on status and ticket information.

  • Drive operational improvements in change management and daily operations by following procedures.

  • Manage and operate large scale IP network technologies and infrastructures.

  • Utilize your skills in Peering and Datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits.

  • Monitor and support the network health of on-premises and cloud infrastructures.

  • Collaborate and develop workflow enhancements while documenting best practices.

What we need to see:

  • Deep knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec.

  • 5+ years of experience in network operations.

  • Skilled in network troubleshooting techniques and demonstrating creative problem-solving abilities.

  • Strong track record of alert response within defined SLAs and Incident management.

  • Experience with one or more of the following CSP environments: AWS, Azure, GCP, OCI.

  • Familiarity with Arista, Fortinet and Juniper.

  • Hands-on experience with contributing to tooling and automation for provisioning, monitoring, and managing complex network infrastructures.

  • Bachelor’s degree in Computer Science, related technical field, or equivalent experience.

  • Excellent verbal and written communication skills.

Ways To Stand Out From The Crowd:

  • Solid understanding of Mellanox/Cumulus OS and Infiniband technology.

  • Skilled in Unix/Linux system administration, with the ability to write and understand Python/Shell scripts to improve efficiency in hyperscale environments.

  • Familiarity with leveraging tools such as Netbox/Nautobot, Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate about innovating and investing in ground breaking technologies.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these outstanding opportunities in deep learning cloud solutions.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 136,000 USD - 224,250 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 2, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Network Reliability Engineer - DGX Cloud in Santa Clara, CA vacancy
  • $152k - $287.5k

     ...NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure...  ...We are looking for Software Engineers with SRE or Production...  ..., GPU systems, CPU systems, networking, Linux, and Kubernetes; turn...  ...Experience managing production reliability through on-call duties, incident... 
    Senior
    Network
    Permanent employment
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $152k - $287.5k

     ...hiring experienced software engineers with kubernetes experience to...  ...: You will be part of an DGX Cloud team responsible for...  ...that enable industry leading reliability, availability, and scalability...  ...diagnostics to cluster and network telemetry. Working with teams... 
    Senior
    Network
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $168k - $322k

     ...As a Senior Software Engineer on NVIDIA’s Global Network Visibility (GNV) team within NVIDIA's Global Network Infrastructure (GNI) organization, you will...  ...roadmaps; make clear tradeoffs among speed, cost, reliability, and maintainability; and mentor engineers through... 
    Senior
    Network
    Full time

    NVIDIA

    Santa Clara, CA
    7 days ago
  • $120k - $171k

     ...have a strong desire to learn about the reliability challenges associated with new product...  ...and Experience PhD in Electrical Engineering, Materials Science, or Physics Strong...  ...across fixed, mobile and transport networks, powered by the innovation of Nokia Bell... 
    Senior
    Network
    Full time
    Temporary work

    Nokia

    Sunnyvale, CA
    2 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide...  ...degradation involving Kubernetes, Linux hosts, storage, network,scheduling, job orchestration, or dependency failures.... 
    Senior
    Network
    Contract work

    VDart

    Santa Clara, CA
    19 hours ago
  • $90k - $130k

     ...Senior Database Reliability Engineer Must Have Technical/Functional Skills • 5+ years of experience designing...  ...with previous work with cloud infrastructure and managed data services...  ...database, operating system, storage, and network layers. • 3+ years of experience in... 
    Senior
    Network
    Remote work

    Tata Consultancy Services

    San Jose, CA
    1 day ago
  • $140k - $270.25k

     ...NVIDIA DGX Cloud provides the infrastructure and software platform that enables enterprises...  ...management is critical to delivering reliable customer experiences while maximizing...  .... We are looking for a Senior Software Engineer to design and build the systems that... 
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $192.4k - $275.8k

     ...the team that keeps Splunk Cloud running for some of the world...  ...enterprise customers, blending Site Reliability Engineering, Systems Engineering, and...  ...You will be the most senior technical individual contributor...  .... Add to that our worldwide network of doers and experts, and... 
    Senior
    Network
    Full time
    Temporary work
    Local area
    Flexible hours

    Cisco

    San Jose, CA
    1 day ago
  • $132.6k - $214.5k

     ...Mission At Palo Alto Networks®, we’re united by a...  ...collaborate closely with our engineering teams to develop...  ...performance and health. As a Senior Staff SRE with the...  ...team, you will: Cloud Expertise: Utilize...  ...and ensure the reliability and availability of our... 
    Senior
    Network
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $104.9k - $174.7k

     ...Site Reliability Engineer The Site Reliability Engineer role is responsible for improving the reliability...  ...to Operations Team projects involving cloud, on-premises infrastructure, security,...  ...support hardware, software, storage, network, cloud, Kubernetes, and other... 
    Senior
    Network
    Temporary work
    Local area

    RELX

    San Jose, CA
    1 day ago
  •  ..., Mirantis empowers platform engineering teams to deliver composable,...  ...environment—on-premises, in the cloud, at the edge, or in sovereign...  ...to GPU architecture, networking, and orchestration. Own the...  ...the reference-system families (DGX, HGX, MGX); aware of what's coming... 
    Senior
    Network

    Mirantis

    San Jose, CA
    8 days ago
  • $183k - $240k

     ...Antora Energy delivers affordable, reliable energy to industry, data centers, and...  ...Antora Energy is seeking an engineer to own the real-time cloud infrastructure connecting our software...  ...and Lambda and manage orchestration, networking, secrets, access controls, and supporting... 
    Senior
    Network
    Remote work
    Flexible hours

    Antora Energy

    San Jose, CA
    12 days ago
  • $184k - $356.5k

     ...NVIDIA is looking for a Senior Software Engineer in Object Storage to design, implement, and extend...  ...users Extending the availability and reliability of our deployments at scale – 10k+...  ...Skilled with building and delivering cloud services, with specific focus on storage... 
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $175k - $229k

     ...DevOps Engineer Instrumental builds the manufacturing acceleration platform...  ...commercial SaaS platforms on public cloud infrastructure, AWS preferred. ~...  ...KPIs to ensure ongoing performance, reliability and efficiency. ~ Network/application security and compliance... 
    Senior
    Network

    Instrumental Inc

    Palo Alto, CA
    2 days ago
  •  ...threats across hybrid multi-cloud environments - stopping the...  ...Team's Vision: Our Engineering team is shaping the future...  ...looking for an experienced Senior Site Reliability Engineer (SRE) with a strong...  ...including compute, storage, networking, and security services ~... 
    Senior
    Network
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    3 days ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable...  ...environments in production and cloud environments. Manage PostgreSQL deployments...  ..., operating system, storage, and network layers. Design and maintain ETL... 
    Senior
    Network
    Full time

    Neshent Technologies

    San Jose, CA
    a month ago
  •  ...operations. Bitdeer also offers advanced cloud capabilities to customers with high demand...  .... Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework...  ...cost-optimizer, gpu-efficiency-dashboard, network-stability-dashboard, patching-... 
    Senior
    Network
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    2 days ago
  • $40 - $55 per hour

     ...55,000 staff across a decentralised and entrepreneurial network of 900 laboratories in over 50 countries. Eurofins offers...  ...products. Job Description Eurofins E&E is seeking a Senior Reliability Test Engineer – Environmental Simulation Lab to perform, document, and... 
    Senior
    Network
    Hourly pay
    Full time
    Contract work
    Work at office
    Monday to Friday

    Eurofins USA Consumer Product Testing

    Santa Clara, CA
    a month ago
  • $160k - $170k

     ...This range is provided by Index Engines. Your actual pay will be based on your skills...  ...Overview Index Engines is seeking a Senior Cloud Engineer with strong expertise in AWS...  ...Familiarity with server hardware and networking. Familiarity with DevOps practices.... 
    Senior
    Network
    Full time

    Index Engines

    San Jose, CA
    3 days ago
  •  ...We Are Synopsys is the leader in engineering solutions from silicon to systems, enabling...  ..., and increase infrastructure reliability across environments that support critical...  ...platforms, compute infrastructure, storage, networking, cloud services, and business-critical... 
    Senior
    Network

    Synopsys

    Sunnyvale, CA
    2 days ago
  •  ...Synopsys is the leader in engineering solutions from silicon to systems...  ...in environments where reliability and speed both matter, and you...  ...prem data centers and public cloud environments (Azure, AWS, and...  ..., including core compute, networking, identity, and security fundamentals... 
    Senior
    Network
    Work at office
    Relocation

    Synopsys

    Sunnyvale, CA
    13 hours ago
  • $208k - $327.75k

     ...lasting impact on the world. We are looking for a Senior Product Manager to define and own the DGX Cloud Data Platform and Analytics System! This platform...  ...the data. It works in close partnership with engineering leadership, infrastructure security, and compliance... 
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...manages the account, access, and network foundation that Lambda's cloud...  ...Security, GRC, IT, and the engineering teams that consume the foundation...  ...you own well-documented, reliable, and observableWho You Are6+... 
    Senior
    Network
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    18 hours ago
  • $133.2k - $192.8k

     ...Job Description: Altera is seeking a Senior Endpoint Engineer specializing in Azure Virtual Desktop...  ...pipelines Proficiency with Azure networking concepts including private endpoints,...  ...Bachelor's Degree in Information Technology, Cloud Computing, or a related field Job... 
    Senior
    Network
    Full time
    Work at office
    Local area
    Remote work
    Shift work

    Altera

    San Jose, CA
    1 day ago
  • $175k - $265k

     ...infrastructure layer that every engineering team and customer depends on...  ..., on-premises GPU clusters, cloud environments, and the...  ...that team, responsible for reliability, automation, and observability...  ...provisioning, OS configuration, networking, storage, and hardware... 
    Senior
    Network
    Full time

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $262k - $364k

     ...behavioral patterns.Build the metrics engines and dashboards to track...  ...scientific concepts into reliable developer tools.Strong understanding...  ..., large-scale system design, networking and data storage, security,...  ...software solutions. Google Cloud is building the most... 
    Senior
    Network

    Google

    Sunnyvale, CA
    3 days ago
  • $95.91k - $144.6k

     ...salary ranges are determined by role, level, qualifications and work location. Client: Power Integrations Senior Reliability Engineer Quality Assurance San Jose, California Description Senior Reliability Engineer Job Description:... 
    Senior
    Work experience placement

    Netpace

    San Jose, CA
    19 hours ago
  • $140k - $215k

     ...services in the CrowdStrike cloud that expose sub-second, near-...  ...CrowdStrike is seeking a Senior Software Developer in Test (SDET...  ...contract, and chaos/resilience engineering. Cybersecurity experience is...  ...of level or role Employee Networks, geographic neighborhood groups... 
    Senior
    Network
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  • $130.6k - $285k

     ...shuffle service, cluster scheduler, and reliability tooling that powers the company's...  ...tooling. About the Role As a Senior Software Engineer on Spark Platform, you will set the...  ...operator patterns, executor pod lifecycle, network topology, and the multi-tenant... 
    Senior
    Network
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    DoorDash USA

    Sunnyvale, CA
    2 days ago
  • $116k - $184k

     ...profoundly impacting society. Come join the team and help build the next era of computing!We're seeking an outstanding Senior HTOL Reliability Engineer to join our Santa Clara lab. This role requires deep device-circuitry knowledge and hands-on hardware development. You... 
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Network Reliability Engineer - DGX Cloud. Be the first to apply!