Senior Site Reliability Engineer, AIOPs
$148k - $235.75kJobleads-US
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job‑centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer to operate the platform itself (not the compute cluster): uptime, performance, data integrity, and safe change management. You’ll own SLOs/SLIs, incident response, and postmortems for the telemetry ingestion, processing, storage, and APIs/dashboards that operators depend on. You’ll partner Software Engineering and Systems Engineering team to translate platform signals into actionable, trustworthy alerts and automation.
What you'll be doing:
- Continuously monitor platform health via dashboards/logs/metrics, automate recurring checks, and keep reliability + resource efficiency on track.
- Own Kubernetes deployments end-to-end (runbooks, canary checks, post‑deploy validation), and lead rollbacks/remediations when needed.
- Lead first‑level incident triage: collect diagnostics, identify likely root causes, and hand off clear, actionable findings to engineering.
- Build and maintain runbooks/SOPs/checklists, pushing continuous improvement through automation.
- Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.
- Contribute in adjacent functional areas to grow and help your team members!
What we need to see:
- BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.
- Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on‑call, addressing incidents, and follow‑up evaluations that drive measurable improvements.
- Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry‑heavy microservices—ingestion, processing, storage, APIs, and UI.
- Automation‑first approach: solid scripting (Python/Bash), CI/CD, and infrastructure‑as‑code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.
- Clear communicator who writes excellent runbooks/docs and can translate ambiguous requirements into concrete operational practices and dependable customer‑facing reliability.
Ways to stand out from the crowd:
- Strong Linux + networking fundamentals, distributed systems instincts, and hands‑on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.
- Experience building safe automation that operators trust: canary releases, automated rollback criteria, “monitoring for the monitoring” (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.
- Strong in distributed/streaming systems operations (Kafka/Pulsar, Flink/Spark, ClickHouse/Elastic/TSDBs, object storage)—and can reason about backpressure, hotspots, and failure domains end‑to‑end.
- Proven programming experience building automation tools or services — ideally in Python, or similar languages — to simplify operations and scale recurring processes.
- Proven experience running large‑scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‑on experience with observability tools — you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.
With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD for Level 3, and 176,000 USD - 276,000 USD for Level 4. You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until May 16, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry.
Learn more about NVIDIA.
#J-18808-Ljbffr Jobleads-US$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...Senior$160k - $240k
...millions of times a day - quickly, reliably, and securely. Any time you... ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our... ...operations or DevOps at a mid-to-senior level.Strong shell scripting...SeniorFull time$104.9k - $174.7k
...SRE role is responsible for improving the reliability, availability, performance, and... ...actions through completion.Follow up with engineering, development, security, support, and business... ...Qualifications5+ years of experience in Site Reliability Engineering, Systems Engineering...SeniorFull timeLocal area$262k - $364k
...infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while working closely with senior technical leads in the development teams.... ...:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software...Senior$132.6k - $214.5k
...you will collaborate closely with our engineering teams to develop innovative solutions that... ...’ performance and health. As a Senior Staff SRE with the Cortex Observability... ...operability of the product and ensure the reliability and availability of our services....SeniorFull timeWork at officeVisa sponsorshipWork visa$150.4k - $277.6k
...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long... ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced...SeniorRelocationDay shift$175k - $265k
...owns the infrastructure layer that every engineering team and customer depends on —... ...core member of that team, responsible for reliability, automation, and observability across colo... ...Splunk, or equivalent), contributing to AIOps-driven detection workflows.Participate in...Senior- ...We Are Synopsys is the leader in engineering solutions from silicon to systems, enabling... ...complexity, and increase infrastructure reliability across environments that support critical... ...software engineering, platform engineering, site reliability engineering, or...Senior
- ...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines... ...the team for you. Your Impact You will be the most senior technical individual contributor on the team — setting the...Senior
- ...Own the architecture and design of reliable, scalable, cost-effective, and performant AI... ...master's degree in computer science or engineering is preferred. Key Skills Software... ...Machine Learning, Artificial Intelligence, Site Reliability Engineering, AI Infrastructure...Senior
$262k - $364k
...Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI Google Sunnyvale, CA, USA X In most instances, this position requires in-person interviews as part of the hiring process. ~ Bachelor’s degree in Computer Science, a related field, or equivalent...Senior$195k - $285k
...infrastructure underpinning our engineering organization must be as reliable and scalable as the chips... ...and leads d-Matrix's Site Reliability Engineering... ..., and serve as the senior escalation point for incidents... ...fleet auto-remediation, and AIOps-driven alerting.Own...Remote work- ...The Role The Platform Engineering team builds, secures, and... ...components deployed at customer sites. The Site Reliability Engineering discipline... ...mentor engineers at different seniority levels, set standards... ...healing workflows powered by AIOps, including Amazon Bedrock...Work at officeRemote work
$122.5k - $175k
...era? Join us at Zscaler. Role We are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San... ...AI-driven anomaly detection, intelligent log parsing, or AIOps tools to automate root-cause analysis and predictive auto-scaling...Full timeWork at officeLocal area3 days per week- ...Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Responsibilities... ...continuous improvement initiatives to enhance reliability, scalability, and efficiency of... ...Bachelor’s degree in computer science, Engineering, or related field; or equivalent work experience...SeniorWork experience placement
$224k - $356.5k
...smart personal assistants and engineering-productivity tools to data-... ...across the company. Now we need a senior staff-level, hands-on... ...for someone who obsesses over reliability, polish, and user trust — and... ...productivity, engineering efficiency, AIOps, and enterprise operations....SeniorFull timeLive in$152k - $241.5k
...infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked... ...initial provisioning through repair.Experience managing production reliability through on-call duties, incident response, observability, and...SeniorPermanent employmentFull time$193.3k - $261.5k
...seeking an experienced Software Development Engineer to build virtualized, hardware-... ...celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring and... ...or architecture (design patterns, reliability and scaling) of new and existing systems...SeniorInternshipLocal areaFlexible hours$248k - $396.75k
...Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems... ..., and automated anomaly detection. Partner with senior leaders and engineers across Cloud, Platform, Security, Networking...$152k - $241.5k
NVIDIA is seeking an innovative and highly motivated engineer with deep expertise in systems software to join our GPU Software team. In... ...environmentsCollaborate with globally distributed teams to deliver scalable, reliable, and high-impact GPU software solutionsWhat we need to see: BS...SeniorFull time$307k - $427k
...with Silicon, Android OS, and Google Research teams to define the long-term roadmap for spatial intelligence.Mentor senior technical leads and staff engineers, fostering a culture of technical excellence and innovative problem-solving across the organization.Influence the...Senior- ...NVIDIA Corporation in Santa Clara, CA is hiring a DevOps Engineer to operate our AI Data Center telemetry platform. You’ll own reliability, incident response, and postmortems for telemetry ingestion, processing, storage, and APIs/dashboards used by operators. Expect...Senior
$207k - $300k
...areas within SU SRE, mentoring team members to enhance system reliability and efficiency.Initiate, own, and lead large-scale,... ...Design for Reliability techniques.3 years of experience as a Site Reliability Engineer.3 years of experience leading projects.3 years of experience...$170k - $200k
...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance...Full time$164k - $205k
Join to apply for the Senior Software Engineer role at Cohesity Join to apply for the Senior Software Engineer role at Cohesity Get AI-powered... ....00-$355,000.00 1 week ago Sr Principal Engineer Software (AIOps for NGFW) Senior Software Engineer, Fabric Networking - GPU...SeniorFull timeWork at officeRemote work2 days per week3 days per week$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...Night shift$145k - $165k
...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining...Work at officeImmediate start$104.4k - $171k
...The mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of production environments, enterprise-grade cloud data reliability, and service continuity for the Cloud Intelligence Group. Our greatest challenge lies in...$230k - $250k
...minds are shaping the future of network reliability, security, and AI‑ready operations. About... ...you will be building the reliability engineering function at Forward — defining how we... ...Looking For ~6+ years of experience in site reliability engineering, DevOps, or...Night shift- ...Investigate and resolve performance and reliability issues across application, infrastructure, database, Kubernetes, and Linux layers... ..., plan capacity, improve observability, and collaborate with engineering teams and business stakeholders. Requirements: Requires hands...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer, AIOPs. Be the first to apply!
- site reliability engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- senior computer engineer Santa Clara, CA
- senior manager customer operations Santa Clara, CA
- senior development engineer Santa Clara, CA
- senior software engineer ruby on rails Santa Clara, CA
- sr marketing manager Santa Clara, CA
- senior customer service Santa Clara, CA
- senior business manager Santa Clara, CA
- senior account executive Santa Clara, CA




