Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer, AIOPs

$148k - $235.75k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer to operate the platform itself (not the compute cluster): uptime, performance, data integrity, and safe change management. You’ll own SLOs/SLIs, incident response, and postmortems for the telemetry ingestion, processing, storage, and APIs/dashboards that operators depend on. You’ll partner Software Engineering and Systems Engineering team to translate platform signals into actionable, trustworthy alerts and automation.**What you'll be doing:*** Continuously monitor platform health via dashboards/logs/metrics, automate recurring checks, and keep reliability + resource efficiency on track.* Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation), and lead rollbacks/remediations when needed.* Lead first-level incident triage: collect diagnostics, identify likely root causes, and hand off clear, actionable findings to engineering.* Build and maintain runbooks/SOPs/checklists, pushing continuous improvement through automation.* Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.* Contribute in adjacent functional areas to grow and help your team members!**What we need to see:****Ways to stand out from the crowd:**With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD for Level 3, and 176,000 USD - 276,000 USD for Level 4.You will also be eligible for equity and .Applications for this job will be accepted at least until May 16, 2026.This posting is for an existing vacancy.NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.* BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.* Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.* Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry-heavy microservices—ingestion, processing, storage, APIs, and UI.* Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.* Clear communicator who writes excellent runbooks/docs and can translate ambiguous requirements into concrete operational practices and dependable customer-facing reliability.* Strong Linux + networking fundamentals, distributed systems instincts, and hands-on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.* Experience building safe automation that operators trust: canary releases, automated rollback criteria, “monitoring for the monitoring” (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.* Strong in distributed/streaming systems operations (Kafka/Pulsar, Flink/Spark, ClickHouse/Elastic/TSDBs, object storage)—and can reason about backpressure, hotspots, and failure domains end-to-end.* Proven programming experience building automation tools or services — ideally in Python, or similar languages — to simplify operations and scale recurring processes.* Proven experience running large‐scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‐on experience with observability tools — you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar. #J-18808-Ljbffr NVIDIA Corporation

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer, AIOPs in Santa Clara, CA vacancy
  • $151.6k - $245.3k

     ...outcomes. Job Summary Palo Alto Networks runs a large hybrid infrastructure and is one of the largest GCP customers. As a Site Reliability Engineer, you will be part of a team supporting the services running on this infrastructure. This includes automation, architecture... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks, Inc.

    Santa Clara, CA
    2 days ago
  • $168k - $270.25k

    Senior Site Reliability Engineer page is loaded## Senior Site Reliability Engineerlocations: US, CA, Santa Claratime type: Full timeposted on: Posted Yesterdayjob requisition id: JR2017460NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    16 hours ago
  • $174k - $252k

    Senior Software Engineer, Site Reliability Engineering X Applicants in San Francisco: Qualified applications with arrest or conviction records will be considered for employment in accordance with the San Francisco Fair Chance Ordinance for Employers and the California... 
    Senior
    Full time

    Google Inc.

    Sunnyvale, CA
    4 days ago
  • $200k - $322k

    Senior Manager, Site Reliability Engineering page is loaded## Senior Manager, Site Reliability Engineeringlocations: US, CA, Santa Claratime type: Full timeposted on: Posted Yesterdayjob requisition id: JR2016119For over 25 years, NVIDIA has been at the forefront of transforming... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    1 day ago
  • $145k - $165k

    A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    1 day ago
  • $176k - $276k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    16 hours ago
  • $180k - $200k

     ...Holmdel, NJ. Join us and be part of a team that's shaping the future of payments—one experience at a time. As our Site Reliability Engineer, you will design, build, and maintain the systems and infrastructure that power our applications, ensuring their... 
    Senior
    For contractors
    Work at office
    Work from home
    Flexible hours

    PayNearMe, Inc.

    Santa Clara, CA
    1 day ago
  • $210k - $270k

    Zocdoc is seeking a Senior Site Reliability Engineer to develop and maintain distributed production systems. The ideal candidate will have over 5 years of experience in site reliability or production engineering, particularly in cloud environments like AWS. Responsibilities... 
    Senior

    GoTo Meeting

    Palo Alto, CA
    4 days ago
  • $207k - $300k

    Google Inc. is looking for a Staff Software Engineer specializing in Site Reliability Engineering in Sunnyvale, CA. This role combines software and systems engineering to build and manage distributed systems, ensuring high reliability and uptime. The ideal candidate should... 
    Senior

    Google Inc.

    Sunnyvale, CA
    1 day ago
  •  ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity...  ...: We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background... 
    Senior
    Work experience placement

    Illumio

    Sunnyvale, CA
    3 days ago
  • $140k - $220k

    About the Job You’ll own reliability and operational excellence for Pylon’s production systems. This means designing and implementing...  ...scale as we grow. You’ll build tooling that makes the entire engineering team more effective, establish on‑call rotations and runbooks... 
    Senior

    Pylon

    Palo Alto, CA
    1 day ago
  • $90k - $140k

    Tata Consultancy Services Limited is looking for a Site Reliability Engineer in Sunnyvale, CA, with 8-10 years of experience in application support across multiple environments. The role involves end-to-end ownership of production environments, ensuring reliability, and... 
    Senior

    Tata Consultancy Services Limited

    Sunnyvale, CA
    4 days ago
  • $210k - $270k

    Your Impact on our Mission: Zocdoc is looking for a Senior Site Reliability Engineer to help develop, monitor, and maintain our distributed production systems. You’ll be challenged with building frameworks and processes for ensuring uptime for our patients and providers... 
    Senior
    Flexible hours

    GoTo Meeting

    Palo Alto, CA
    4 days ago
  • $174k - $252k

    A leading tech company is seeking a Senior Software Engineer for Site Reliability Engineering based in Sunnyvale, CA. The role involves ensuring service reliability, leading technical projects, and enhancing systems performance. Candidates should have at least 5 years of... 
    Senior

    Google Inc.

    Sunnyvale, CA
    4 days ago
  • $80 per hour

     ...Senior Cloud DevOps Engineer/Site Reliability Engineer Position Title: Senior Cloud DevOps Engineer/Site Reliability Engineer Location: San Jose, CA (Look for local profiles only) Duration: 6 Months Bill Rate: $80/hr all inclusive (First priority for W2 but... 
    Senior
    Local area

    ClifyX

    San Jose, CA
    1 day ago
  • A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 
    Senior

    TechDigital Group

    Santa Clara, CA
    16 hours ago
  • $180k - $260k

     ...facilitating effortless integration into customers’ logistics operations. About the role We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring, and scaling of our growing fleet of autonomous vehicles. In this role, you... 
    Senior
    Odd job
    Work at office
    Remote work

    Booster

    Mountain View, CA
    4 days ago
  • A leading technology company is looking for a Java SRE Engineer to support large-scale cloud migrations and production systems on AWS...  ...mentoring team members and collaborating with various teams to ensure reliability. This position is onsite in the San Francisco Bay Area. #J-188... 
    Senior

    EITACIES Inc.

    Santa Clara, CA
    4 days ago
  • A leading tech recruiting firm is seeking a Site Reliability Engineer to manage and optimize cloud infrastructure primarily using GCP or AWS. The role involves maintaining high availability through Kubernetes clusters and improving CI/CD pipelines with Terraform. Ideal... 
    Senior

    Amiri Recruiting

    Mountain View, CA
    4 days ago
  • $147k - $237.5k

     ...Knowledge of Linux fundamentals and networked computing environment concepts. Additional Information: The Team: Our engineering team is at the core of our products and connected directly to the mission of preventing cyberattacks. We are constantly innovating... 
    Full time
    Work at office
    Local area

    Palo Alto Networks

    Santa Clara, CA
    5 days ago
  • $120.3k - $194.53k

     ...drives great outcomes. Job Summary Palo Alto Networks runs a large hybrid infrastructure across multiple public clouds. As a Site Reliability Engineer on the Internet Security Platform team, you will be part of a team supporting Advanced DNS Security services. This... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks, Inc.

    Santa Clara, CA
    16 hours ago
  • $147k - $237.5k

    Palo Alto Networks, Inc. is seeking a skilled software engineer with over 5 years of experience in building enterprise applications. This role emphasizes expertise in Java programming and working with distributed systems. The position involves designing advanced data processing... 

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...autonomous vehicles. We are now looking for a Senior Software Engineer to help accelerate the next era of...  ...to ensure delivery of functional, reliable, secure, and performance-optimal GPU clusters...  ...You will also research in traditional AIOps and the emerging Agentic AI, and... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  •  ...Job Description Job Description Forhyre is looking for engineers who can bring unique perspectives and innovative ideas to all areas...  ...evangelize cloud best practices while building a culture of reliability and observability Engage in and improve the end to end lifecycle... 

    Forhyre

    Sunnyvale, CA
    5 days ago
  •  ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the... 
    Work at office

    Foxconn Industrial Internet - FII

    Sunnyvale, CA
    26 days ago
  •  ...of Huobi globe spanning infrastructure. •       Work with engineering teams to make sure new features and changes are deployed quickly...  .... •       Constantly improve our system performance and reliability through better tools, process and monitoring system. •... 
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    5 days ago
  • $202k - $247k

    Job Category Site Reliability Engineering Posting Date 11/18/2025, 12:24 AM Locations Santa Clara, CA, United States Job Schedule Full time Job Description...  ...of DevOps/SRE experience, with at least 5 years in a senior or lead role managing production systems at scale. Expert-... 
    Full time
    Worldwide

    Fortinet, Inc.

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...We are now looking for a Systems Software Engineer. Do you like to think creatively and enjoy solving challenges that require innovation? If so, we may have an opportunity for you. In Our team we define and build methodologies, Software, and flows tailored to the field... 
    Senior

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...NVIDIA is searching for a creative and highly motivated engineer with expertise in system s software to join the GPU Software team. You will design key aspects of our production GPU kernel drivers and embedded SW that impacts our products both in the datacenter and in... 
    Senior

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...Senior Release Engineer, Hyperscale Line Of Business Santa Clara, California We're in an unbelievably exciting area of tech and are fundamentally reshaping the data storage industry. Here, you lead with innovative thinking, grow along with us, and join the smartest... 
    Senior
    Work at office
    Flexible hours

    Pure Storage

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer, AIOPs. Be the first to apply!