Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Network Reliability Engineer - DGX Cloud

$136k - $224.25k

2100 NVIDIA USA

NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence. In this role, the Senior Network Operations Engineer will remediate critical alerts within defined SLAs, triage production impacting network incidents, and interact with internal customers on network related issues. They will also be responsible for engaging with external vendors to remediate hardware and software issues, and participate in project related work such as network device upgrades and capacity augmentations. An ideal candidate will possess a wide range of skills, including alert monitoring & resolution in large-scale networks and CSP environments, outstanding troubleshooting skills, understanding of L3 underlay networks, and network protocol knowledge in large multi-vendor infrastructures. What you will be doing: Engage in 24/7 global shift rotations to provide remote support for network repairs and changes while collaborating across teams and updating customers on status and ticket information. Drive operational improvements in change management and daily operations by following procedures. Manage and operate large scale IP network technologies and infrastructures. Utilize your skills in Peering and Datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits. Monitor and support the network health of on-premises and cloud infrastructures. Collaborate and develop workflow enhancements while documenting best practices. What we need to see: Deep knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec. 5+ years of experience in network operations. Skilled in network troubleshooting techniques and demonstrating creative problem-solving abilities. Strong track record of alert response within defined SLAs and Incident management. Experience with one or more of the following CSP environments: AWS, Azure, GCP, OCI. Familiarity with Arista, Fortinet and Juniper. Hands-on experience with contributing to tooling and automation for provisioning, monitoring, and managing complex network infrastructures. Bachelor’s degree in Computer Science, related technical field, or equivalent experience. Excellent verbal and written communication skills. Ways To Stand Out From The Crowd: Solid understanding of Mellanox/Cumulus OS and Infiniband technology. Skilled in Unix/Linux system administration, with the ability to write and understand Python/Shell scripts to improve efficiency in hyperscale environments. Familiarity with leveraging tools such as Netbox/Nautobot, Prometheus, Grafana, Panoptes to monitor and manage a global network. Passionate about innovating and investing in ground breaking technologies. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you. NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these outstanding opportunities in deep learning cloud solutions. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 136,000 USD - 224,250 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until May 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr 2100 NVIDIA USA

Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Senior Network Reliability Engineer - DGX Cloud in Santa Clara, CA vacancy
  • 2100 NVIDIA USA is seeking a Senior Network Reliability Engineer to maintain their cloud and datacenter network infrastructures. This role includes critical alert remediation, incident triage, and vendor engagement. Candidates should have strong TCP/IP knowledge and 5+... 
    Senior
    Network

    2100 NVIDIA USA

    Santa Clara, CA
    14 hours ago
  • $136k - $224.25k

    NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence... 
    Senior
    Network
    Remote work
    Shift work

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $152k - $241.5k

     ...outstanding, passionate, and dedicated Senior AI Infrastructure Engineer to join our DGX Cloud group. This engineering role...  ...across different systems, networking, coding, databases, capacity management...  ...cloud services deliver maximum reliability and uptime. They carefully... 
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $184k - $287.5k

    The DGX Cloud organization at NVIDIA brings together cutting...  ...a team of innovative engineers dedicated to solving...  ...for an outstanding Senior Systems Software Engineer...  ...such as GPU Operator, Network Operator, DCGM, NIM,...  ...large scale, ensuring reliability and efficiency. Build... 
    Senior
    Network
    Worldwide

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $224k - $356.5k

     ...the world. As part of the DGX Cloud organization, the Attestation...  ...security, silicon, and cloud engineering teams to turn embedded...  ...attestation standards into reliable, self‑service cloud capabilities...  ...across Data Center, Automotive, Networking, and AI ecosystems. Define... 
    Senior
    Network
    Remote work

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $166k - $244k

    Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run...  ...systems. SRE ensures that Google Cloud's services—both our internally critical...  ...apart so we can rebuild them. We keep our networks up and running, ensuring our users... 
    Senior
    Network
    Full time

    Google

    Sunnyvale, CA
    14 hours ago
  • NVIDIA is hiring engineers to scale up its AI Infrastructure. We expect...  ...hardware, operations, and networking, familiarity with software...  ...lifecycle management across cloud providers. Implement...  ...that enable industry-leading reliability, availability, and scalability... 
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $184k - $287.5k

     ...model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage,...  ...art LLM workloads run efficiently and reliably at scale. You will lead deep performance...  ...performance across compute, memory, networking, and communication layers using tools... 
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based...  ...environments. We are looking for a hands‑on senior engineer to own the lifecycle and automation of the Kubernetes... 
    Senior
    Network
    Weekend work

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $156k - $190k

     ...in Sunnyvale, CA, is seeking a Staff Cloud Support Engineer to provide technical leadership in cloud...  ...will lead incident responses, design reliability architecture, and mentor team members...  ...expertise in Linux, Kubernetes, and networking. We offer a competitive salary range... 
    Senior
    Network

    Crusoe Energy Systems

    Sunnyvale, CA
    14 hours ago
  • Production engineering is a field that involves crafting, building, and...  ..., data management, systems, networking, coding, database management,...  ...deployment, along with open‑source cloud‑enabling technologies such as...  ...storage architectures are reliable, scalable, and efficient.... 
    Senior
    Network
    Flexible hours

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • NVIDIA is seeking a Senior AI Infrastructure Engineer to design, build, and maintain large-scale production systems for the DGX Cloud group in California. The role covers software and systems...  ...engineering practices across Linux, networks, storage, and cloud-native tooling... 
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    14 hours ago
  •  ...have a strong desire to learn about the reliability challenges associated with new product...  ...Skills And Experience PhD in Electrical Engineering, Materials Science, or Physics Strong...  ...expertise across fixed, mobile and transport networks, powered by the innovation of Nokia... 
    Senior
    Network
    Temporary work

    Nokia

    Sunnyvale, CA
    4 days ago
  • $133.2k - $219.6k

    A leading cloud services provider in Sunnyvale, California is seeking an experienced Operations and Maintenance Specialist. The ideal candidate must be fluent in Chinese, have over 3 years of relevant experience, and capabilities in architecture and performance optimization... 
    Senior
    Network

    Alibaba Cloud

    Sunnyvale, CA
    14 hours ago
  • $272k - $431.25k

    NVIDIA DGX Cloud is scaling GPU infrastructure across...  ...for Principal Software Engineers to help shape the technical...  ..., automation, and reliability across large-scale GPU...  .... This role is for senior technical leaders who...  ...infrastructure, storage, networking, security, and... 
    Network

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • Palo Alto Networks is seeking a Senior Staff Data Center & OpenShift Operations Engineer to ensure high availability of our infrastructure. The role involves monitoring systems, implementing disaster recovery strategies, and managing OpenShift clusters. Qualified candidates... 
    Senior
    Network

    Palo Alto Networks

    Santa Clara, CA
    14 hours ago
  • A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 
    Senior
    Network

    TechDigital Group

    Santa Clara, CA
    3 days ago
  • $159k - $255k

     ...cybersecurity company in Santa Clara is seeking a Technical Marketing Engineer to develop marketing strategies, create technical tools, and...  ...and sales. Compensation is competitive, with a salary range between $159,000 and $255,000 annually. #J-18808-Ljbffr Palo Alto Networks
    Senior
    Network

    Palo Alto Networks

    Santa Clara, CA
    14 hours ago
  •  ...located in Sunnyvale, California, is seeking a Principal Engineer to lead the reliability and scalability of its cloud infrastructure. The role involves ensuring operational excellence across compute, storage, and networking systems, while mentoring engineers and driving... 
    Senior
    Network

    Crusoe Energy Systems

    Sunnyvale, CA
    14 hours ago
  • Apple Inc. in Cupertino seeks an experienced SRE software engineer to build and enhance compute infrastructure, ensuring services scale to meet demand for Apple Services. You will work on VM lifecycle management and bare metal provisioning within a highly distributed environment... 
    Senior

    Apple Inc.

    Cupertino, CA
    2 days ago
  •  ...Senior Site Reliability Engineer LeanData helps the world's fastest-growing companies automate, simplify...  ...lead the strategic evolution of our cloud infrastructure. Reporting directly to...  ...predictably. Cloud Security: Harden our network architecture and application security... 
    Senior
    Network
    Full time
    Work at office
    Flexible hours
    2 days per week

    LeanData

    Santa Clara, CA
    4 days ago
  • Google in Sunnyvale is seeking an experienced Network Engineer to drive the design and management of complex network infrastructures...  ...enhancing network architectures for Google Distributed Cloud Labs, ensuring reliability and performance at scale. The ideal candidate will have... 
    Senior
    Network

    Google

    Sunnyvale, CA
    1 day ago
  • About the Role Senior Site Reliability Engineer (Payments Infrastructure) - Kody is seeking a Senior Site...  ...response, service-level management, and cloud infrastructure reliability across...  ..., PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms... 
    Senior
    Network

    Kody

    Sunnyvale, CA
    14 hours ago
  • $141.91k - $232.19k

     ...photonics manufacturing and reliability. Responsibilities Assess intrinsic...  ...proper quality systems, engineering shipment management, fab...  ...verbal and written). Strong networking, teamwork, and social skills...  ...concise key messages for senior management review. Influencing... 
    Senior
    Network
    Local area
    Shift work

    Intel

    Santa Clara, CA
    14 hours ago
  • $144k - $209k

    Senior Hardware Reliability Engineer, Global Hardware Reliability Engineering Experience driving progress, solving problems, and mentoring more junior...  ...hardware reliability of new machine learning, server, networking, and storage products. You will also perform early... 
    Senior
    Network
    Contract work

    Google

    Sunnyvale, CA
    14 hours ago
  • $174k - $253k

    Google is seeking a Software Engineer to develop next-generation technologies and work on projects critical to its needs. The role includes...  ...software for TPU supercomputers and implementing efficient network solutions. Ideal candidates will have extensive experience in software... 
    Senior
    Network

    Google

    Sunnyvale, CA
    14 hours ago
  • Crusoe Energy Systems in California is seeking a Staff/Sr. Staff+ Network Engineer to design and optimize our global network infrastructure. The role involves collaboration with cross-functional teams and external vendors to ensure effective network operations and strategies... 
    Senior
    Network

    Crusoe Energy Systems

    Sunnyvale, CA
    14 hours ago
  • $180k - $260k

     ...We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring...  ...manage rollouts of both on-premises and cloud infrastructure in support of...  ...Engineer. ~ Strong knowledge of networking fundamentals, including protocols,... 
    Senior
    Network
    Odd job
    Work at office
    Remote work

    Gatik AI

    Santa Clara, CA
    2 days ago
  • Oracle is seeking experienced Linux Kernel Developers to advance the Linux operating system for large-scale cloud environments. This role involves contributing to the Linux kernel and collaborating on projects across various subsystems. Candidates should have several years... 
    Senior
    Network

    Oracle

    Santa Clara, CA
    1 day ago
  •  ...cybersecurity firm seeks a Sr Staff Software Engineer to design and build scalable distributed backend services for their innovative cloud security solutions. The ideal candidate...  ..., fostering collaboration among innovative teams. #J-18808-Ljbffr Palo Alto Networks
    Senior
    Network

    Palo Alto Networks

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Network Reliability Engineer - DGX Cloud. Be the first to apply!