Senior Site Reliability Engineer, AIOPs
NVIDIA Gruppe
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self‑driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIA, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Role Overview You will be building an AI Data Center AIOps platform that turns raw, high‑volume telemetry into reliable, job‑centric insights and automation for GPU fleets. Join our team of innovative engineers who are building this platform and operating it (not the compute cluster): uptime, performance, data integrity, and safe change management. You’ll own SLOs/SLIs, incident response, and postmortems for the telemetry ingestion, processing, storage, and APIs/dashboards that operators depend on. You’ll partner with the Software Engineering and Systems Engineering team to translate platform signals into actionable, trustworthy alerts and automation. Responsibilities Continuously monitor platform health via dashboards, logs, and metrics, automate recurring checks, and keep reliability and resource efficiency on track. Own Kubernetes deployments end‑to‑end (runbooks, canary checks, post‑deploy validation), and lead rollbacks/remediations when needed. Lead first‑level incident triage: collect diagnostics, identify likely root causes, and hand off clear, actionable findings to engineering. Build and maintain runbooks, SOPs, and checklists, pushing continuous improvement through automation. Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible. Contribute in adjacent functional areas to grow and help your team members. Qualifications BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops. Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on‑call, addressing incidents, and follow‑up evaluations that drive measurable improvements. Deep Kubernetes and containers experience (deploying, debugging, scaling) for telemetry‑heavy microservices—ingestion, processing, storage, APIs, and UI. Automation‑first approach: solid scripting (Python/Bash), CI/CD, and infrastructure‑as‑code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil. Clear communicator who writes excellent runbooks/docs and can translate ambiguous requirements into concrete operational practices and dependable customer‑facing reliability. Ways to Stand Out Strong Linux and networking fundamentals, distributed systems instincts, and hands‑on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale. Experience building safe automation that operators trust: canary releases, automated rollback criteria, “monitoring for the monitoring” (lag/drop/error budgets), and replay/backfill pipelines with correctness checks. Strong in distributed/streaming systems operations (Kafka/Pulsar, Flink/Spark, ClickHouse/Elastic/TSDBs, object storage)—and can reason about backpressure, hotspots, and failure domains end‑to‑end. Proven programming experience building automation tools or services—ideally in Python, or similar languages—to simplify operations and scale recurring processes. Proven experience running large‑scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‑on experience with observability tools—you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar. Benefits Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000USD – 235,750USD for Level3, and 176,000USD – 276,000USD for Level4. You will also be eligible for equity and benefits. NVIDIA offers a competitive salaries and generous benefits package. EEO Statement NVIDIA is committed to fostering a diverse work environment and prides itself as an equal opportunity employer. We do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr NVIDIA Gruppe
$152k - $241.5k
...Overview We’re looking for a Senior SRE to join our Compute Farm... ...lifecycle management, fleet reliability/auto‑healing, E2E observability... ...observability or data‑driven operations (AIOps/ML‑driven signals) that... ..., or Ruby. Mentored other engineers and influenced technical...Senior$151.6k - $245.3k
...outcomes. Job Summary Palo Alto Networks runs a large hybrid infrastructure and is one of the largest GCP customers. As a Site Reliability Engineer, you will be part of a team supporting the services running on this infrastructure. This includes automation, architecture...SuggestedFull timeWork at officeVisa sponsorshipWork visa$200k - $322k
...supportive environment, where NVIDIANs are inspired to excel and make a profound global impact. NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management...Senior$174k - $252k
Senior Software Engineer, Site Reliability Engineering X Applicants in San Francisco: Qualified applications with arrest or conviction records will be considered for employment in accordance with the San Francisco Fair Chance Ordinance for Employers and the California...SeniorFull time- ...cloud‑native infrastructure, where reliability, scale, and intelligent automation define... ...the future of operations. As a Senior Site Reliability Engineer, you will design and operate the platforms... ...tools (e.g., LLM‑based agents, AIOps platforms) into SRE workflows for intelligent...SeniorFull timeWork at officeVisa sponsorshipWork visa
$126k - $204.5k
...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and... ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications Required...Senior$145k - $165k
A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...Senior$176k - $276k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized...Senior- NVIDIA Gruppe in Santa Clara is seeking an experienced engineer to build an AI Data Center AIOps platform. The ideal candidate will have a strong background in Kubernetes and automation, ensuring the reliability of GPU fleet management. Key responsibilities include monitoring...Senior
$210k - $270k
Zocdoc is seeking a Senior Site Reliability Engineer to develop and maintain distributed production systems. The ideal candidate will have over 5 years of experience in site reliability or production engineering, particularly in cloud environments like AWS. Responsibilities...Senior$207k - $300k
Google Inc. is looking for a Staff Software Engineer specializing in Site Reliability Engineering in Sunnyvale, CA. This role combines software and systems engineering to build and manage distributed systems, ensuring high reliability and uptime. The ideal candidate should...Senior$90k - $140k
Tata Consultancy Services Limited is looking for a Site Reliability Engineer in Sunnyvale, CA, with 8-10 years of experience in application support across multiple environments. The role involves end-to-end ownership of production environments, ensuring reliability, and...Senior- The Role We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast-growing...Senior
$140k - $220k
About the Job You’ll own reliability and operational excellence for Pylon’s production systems. This means designing and implementing... ...scale as we grow. You’ll build tooling that makes the entire engineering team more effective, establish on‑call rotations and runbooks...Senior- ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity... ...: We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background...SeniorWork experience placement
$210k - $270k
Your Impact on our Mission: Zocdoc is looking for a Senior Site Reliability Engineer to help develop, monitor, and maintain our distributed production systems. You’ll be challenged with building frameworks and processes for ensuring uptime for our patients and providers...SeniorFlexible hours$174k - $252k
A leading tech company is seeking a Senior Software Engineer for Site Reliability Engineering based in Sunnyvale, CA. The role involves ensuring service reliability, leading technical projects, and enhancing systems performance. Candidates should have at least 5 years of...Senior- A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...Senior
- ...Infrastructure Footprint: Global production infrastructure across AWS, South America, and Europe Role Overview Seeking a Senior Site Reliability Engineer / DevOps Engineer to design, scale, and operate highly available global infrastructure supporting production systems...Senior
$180k - $260k
...facilitating effortless integration into customers’ logistics operations. About the role We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring, and scaling of our growing fleet of autonomous vehicles. In this role, you...SeniorOdd jobWork at officeRemote work- A leading technology company is looking for a Java SRE Engineer to support large-scale cloud migrations and production systems on AWS... ...mentoring team members and collaborating with various teams to ensure reliability. This position is onsite in the San Francisco Bay Area. #J-188...Senior
- Zocdoc, located in Silicon Valley, CA, is seeking a Senior Site Reliability Engineer to monitor and maintain cloud-based systems ensuring uptime for millions of patients. You'll work with cutting-edge technology in a diverse and collaborative environment. This role requires...Senior
- A leading tech recruiting firm is seeking a Site Reliability Engineer to manage and optimize cloud infrastructure primarily using GCP or AWS. The role involves maintaining high availability through Kubernetes clusters and improving CI/CD pipelines with Terraform. Ideal...Senior
$175.8k - $264.2k
Senior Site Reliability Engineer - Apple Services Engineering (ASE) / iCloud Cupertino, CA People at Apple don't just build products - they craft experiences our customers love and depend on. Apple Services Engineering (ASE) builds and supports the systems that make many...Senior- ..., and the challenges of building in a high-growth startup, we’d love to talk. This is more than a job—it’s a journey. Site Reliability Engineers (SREs) are responsible for the overall performance and reliability of ASAPP's infrastructure and products. The team owns...SeniorRemote work
$147k - $237.5k
...Knowledge of Linux fundamentals and networked computing environment concepts. Additional Information: The Team: Our engineering team is at the core of our products and connected directly to the mission of preventing cyberattacks. We are constantly innovating...Full timeWork at officeLocal area$180k - $200k
...Holmdel, NJ. Join us and be part of a team that's shaping the future of payments—one experience at a time. As our Site Reliability Engineer, you will design, build, and maintain the systems and infrastructure that power our applications, ensuring their...SeniorFor contractorsWork at officeWork from homeFlexible hours$120.3k - $194.53k
...drives great outcomes. Job Summary Palo Alto Networks runs a large hybrid infrastructure across multiple public clouds. As a Site Reliability Engineer on the Internet Security Platform team, you will be part of a team supporting Advanced DNS Security services. This...SeniorFull timeWork at officeVisa sponsorshipWork visa$147k - $237.5k
Palo Alto Networks, Inc. is seeking a skilled software engineer with over 5 years of experience in building enterprise applications. This role emphasizes expertise in Java programming and working with distributed systems. The position involves designing advanced data processing...$152k - $241.5k
...autonomous vehicles. We are now looking for a Senior Software Engineer to help accelerate the next era of... ...to ensure delivery of functional, reliable, secure, and performance-optimal GPU clusters... ...You will also research in traditional AIOps and the emerging Agentic AI, and...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer, AIOPs. Be the first to apply!
- site reliability engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- senior game producer Santa Clara, CA
- senior manager process engineering Santa Clara, CA
- senior manufacturing engineer Santa Clara, CA
- senior manager clinical operations Santa Clara, CA
- senior optical engineer Santa Clara, CA
- senior lead project manager Santa Clara, CA
- senior manager quality engineering Santa Clara, CA
- senior device engineer Santa Clara, CA


