Senior Site Reliability Engineer

$148k - $235.75k

NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

NVIDIA is looking for a seasoned SRE to join its complex and fast-paced Infrastructure, Planning and Processes organization where you will be working as a Senior SRE Engineer. The position will be part of a fast-paced crew that develops and maintains sophisticated NVIDIA's internal Jenkins based CI/CD product for GPUs and Tegra systems. The team works with various other business units within NVIDIA Software such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Driverless Cars to cater to their infrastructure & systems needs. As an SRE, you’ll also be working in conjunction with various teams such as software engineering to deploy these new products and handle our infrastructure, associated processes and systems. Keen attention to detail, problem-solving abilities, and a solid knowledge base are needed.

What you’ll be doing:

Manage NVIDIA's on-prem infrastructure. Maintain uptime, reliability and readiness of on-prem engineering cloud spread across multiple data centers.
Guard service level agreements (SLAs) for critical engineering services. Implement monitoring, alerting, and incident response procedures to ensure alignment to defined performance targets. Perform root cause analysis and post-mortems of incidents for any threshold breaches.
Deploy, configure, and manage applications and services on Kubernetes clusters. Implement logging, monitoring, and alerting solutions (e.g., Prometheus, Grafana, ELK/EFK). Ensure high availability, fault tolerance, and disaster recovery for Kubernetes workloads.
Help in capacity planning, optimization and better utilization efforts.
Support user reported issues & issues. Monitor alerts and take necessary action. Actively participate in WAR room for critical issues
Drive automation of monitoring to gain more insight into applications and system health.
Reuse AI techniques to extract useful signals about machines and jobs from the data generated.

What we need to see:

Experience of maintaining cloud infrastructure and highly-available production environment.
Experience handling and maintaining systems installed in on-premises data centers, with strong hands-on proficiency using BMC interfaces (Redfish), KVM, and IPMI tools for hardware provisioning, remote access, and troubleshooting. Knowledge and understanding of Openstack architecture and services is a plus.
Proven background working with databases, including relational databases such as SQL/MySQL, as well as time-series databases like Prometheus, with experience in data querying & performance tuning.
Solid understanding of networking principles and protocols, including TCP/IP, DNS, DHCP, and VLANs, with the ability to diagnose connectivity issues and support complex, distributed systems.
Practical experience in working with data analytics and visualization tools such as Kibana, Grafana, Splunk, or similar platforms, applied to analyze logs, metrics, and system behavior for monitoring and troubleshooting purposes.
Strong demonstrable experience in automation tools like Jenkins and/or Temporal along with configuration tools like Ansible.
Proficiency with Kubernetes, Docker, and virtualization technologies, with experience deploying, managing, and operating containerized workloads and virtualized infrastructure in production environments.
Advanced knowledge of standard security methodologies and protocols, including system hardening, access control, vulnerability management, and secure operations across infrastructure and application layers.
5+ years of demonstrable experience.
Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent experience.

Ways to stand out from the crowd:

Previous experience with SRE teams managing on-prem infrastructure.
Experience managing NVIDIA hardware like GPUs and Tegras.
Thrives in a multi-tasking environment with constantly evolving priorities.
Outstanding interpersonal skills and communication with all levels of management.

With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD for Level 3, and 176,000 USD - 276,000 USD for Level 4.

You will also be eligible for equity and benefits ( .

Applications for this job will be accepted at least until June 13, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Apply

Vacancy posted 4 days ago

Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Santa Clara, CA vacancy

Senior Site Reliability Engineer
$181.69k - $213.75k
...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects founders, investors, and limited partners through world-class software, purpose-built for everyone in venture capital, private...
Senior
Full time
Work at office
Carta
Santa Clara, CA
3 days ago
Senior Site Reliability Engineer
...Senior Site Reliability Engineer LeanData helps the world's fastest-growing companies automate, simplify, and accelerate revenue. We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly...
Senior
Full time
Work at office
Flexible hours
2 days per week
LeanData
Santa Clara, CA
3 days ago
Senior Staff Site Reliability Engineer
$126k - $204.5k
...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and... ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications...
Senior
Full time
Work at office
Palo Alto Networks
Santa Clara, CA
2 days ago
Senior Site Reliability Engineer, AIOPs
...building an AI Data Center AIOps platform that turns raw, high‑volume telemetry into reliable, job‑centric insights and automation for GPU fleets. Join our team of innovative engineers who are building this platform and operating it (not the compute cluster): uptime, performance...
Senior
NVIDIA Gruppe
Santa Clara, CA
1 day ago
Senior Software Engineer, Site Reliability Engineering
$174k - $252k
Senior Software Engineer, Site Reliability Engineering X Applicants in San Francisco: Qualified applications with arrest or conviction records will be considered for employment in accordance with the San Francisco Fair Chance Ordinance for Employers and the California...
Senior
Full time
Google Inc.
Sunnyvale, CA
1 day ago
Senior Site Reliability Engineer - HPC
$152k - $241.5k
...intelligence. Job Overview We’re looking for a Senior SRE to join our Compute Farm team and... ...host lifecycle management, fleet reliability/auto‑healing, E2E observability or data‑... ...Python, Go, Perl, or Ruby. Mentored other engineers and influenced technical direction through...
Senior
NVIDIA Gruppe
Santa Clara, CA
4 days ago
Senior Site Reliability Engineer - HPC
$152k - $241.5k
Overview NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join our Compute Farm team and help build the next generation of our global services platform. The role focuses on keeping critical systems operational while leveraging AI technologies to deliver...
Senior
NVIDIA Corporation
Santa Clara, CA
4 days ago
Senior Manager, Site Reliability Engineering
$200k - $322k
Senior Manager, Site Reliability Engineering page is loaded## Senior Manager, Site Reliability Engineeringlocations: US, CA, Santa Claratime type: Full timeposted on: Posted Yesterdayjob requisition id: JR2016119For over 25 years, NVIDIA has been at the forefront of transforming...
Senior
NVIDIA Corporation
Santa Clara, CA
3 days ago
Senior Site Reliability Engineer — Scale, Automation & Uptime
$145k - $165k
A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...
Senior
Bolt Graphics, Inc.
Sunnyvale, CA
3 days ago
Senior Site Reliability Engineer
...The Role We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast...
Senior
XRC Ventures
Palo Alto, CA
5 days ago
Senior Site Reliability Engineer (SRE)
$158k - $225k
...Senior Site Reliability Engineer (SRE) Manufacturing advanced electronics requires understanding millions of signals generated across complex assembly processes. Instrumental builds systems that capture and analyze those signals — images, test results, and process...
Senior
Instrumental Inc
Palo Alto, CA
2 days ago
Sr. Site Reliability Engineer
...the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity... ...We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background...
Senior
Work experience placement
Immediate start
Illumio
Sunnyvale, CA
5 days ago
Senior Site Reliability Engineer
$150k - $175k
...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed...
Senior
Remote work
ASAPP
Mountain View, CA
4 days ago
Senior Site Reliability Engineer
...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software...
Senior
Immediate start
Remote work
Worldwide
OutSystems
Menlo Park, CA
3 days ago
Senior Site Reliability Engineer
$137.77k - $194.59k
...distributed team of roughly 80 scientists and engineers building and operating Rubin's petascale... .... Your role: You will own the reliability and robustness of Rubin Observatory's... ...of this position, SLAC is open to on-site, hybrid, and remote work options. Work...
Senior
Remote work
Flexible hours
Night shift
Stanford University
Menlo Park, CA
3 days ago
Senior Site Reliability Engineer
$159.2k - $301.6k
...The Opportunity We are seeking a Senior SRE (Site Reliability Engineer) to help compose, build, and operate highly scalable, secure, and resilient cloud platforms. We are redefining this role to focus on product and platform engineering. This position is a core builder...
Senior
Temporary work
Local area
Worldwide
Adobe
San Jose, CA
1 day ago
Senior Site Reliability Engineer
...Senior Site Reliability Engineer Latitude AI develops automated driving technologies, including L3, for Ford vehicles at scale. We're driven by the opportunity to reimagine what it's like to drive and make travel safer, less stressful, and more enjoyable for everyone...
Senior
Work at office
Immediate start
Latitude AI
Palo Alto, CA
3 days ago
Senior Site Reliability Engineer - Remote & Scalable Impact
...join our small team focused on growth and productivity. The role involves scaling our platform and infrastructure while enhancing reliability and the overall developer experience. Ideal candidates will have strong expertise in distributed systems, cloud-native...
Senior
Remote job
BuildBuddy
Palo Alto, CA
1 day ago
Senior SRE Software Engineer - Reliability & Scale
$207k - $300k
Google Inc. is looking for a Staff Software Engineer specializing in Site Reliability Engineering in Sunnyvale, CA. This role combines software and systems engineering to build and manage distributed systems, ensuring high reliability and uptime. The ideal candidate should...
Senior
Google Inc.
Sunnyvale, CA
3 days ago
Senior/Staff Site Reliability Engineer
$180k - $260k
...effortless integration into customers' logistics operations. About the role We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring, and scaling of our growing fleet of autonomous vehicles. In this role, you will...
Senior
Odd job
Work at office
Remote work
Gatik AI
Mountain View, CA
3 days ago
Senior Site Reliability Engineer
$140k - $220k
About the Job You’ll own reliability and operational excellence for Pylon’s production systems. This means designing and implementing... ...scale as we grow. You’ll build tooling that makes the entire engineering team more effective, establish on‑call rotations and runbooks...
Senior
Pylon
Palo Alto, CA
3 days ago
Sr. Site Reliability Engineer
...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity... ...: We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background...
Senior
Work experience placement
Illumio
Sunnyvale, CA
5 days ago
Senior Site Reliability Engineer (SRE)
$181k - $197k
Senior SRE Palo Alto, CA • Engineering • Hybrid • Full-time Founded by a team of ex-Apple engineers, Instrumental provides a collection of software... ...on, and measuring KPIs to ensure ongoing performance, reliability and efficiency. Network/application security and...
Senior
Full time
Clutch Canada
Palo Alto, CA
5 days ago
Senior SRE Engineer: Scale & Reliability Leader
$174k - $252k
A leading tech company is seeking a Senior Software Engineer for Site Reliability Engineering based in Sunnyvale, CA. The role involves ensuring service reliability, leading technical projects, and enhancing systems performance. Candidates should have at least 5 years of...
Senior
Google Inc.
Sunnyvale, CA
1 day ago
Senior Wireless Network SRE & Reliability Engineer
A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless...
Senior
TechDigital Group
Santa Clara, CA
2 days ago
Senior Java SRE & Platform Engineer - AWS/Kubernetes
A leading technology company is looking for a Java SRE Engineer to support large-scale cloud migrations and production systems on AWS... ...mentoring team members and collaborating with various teams to ensure reliability. This position is onsite in the San Francisco Bay Area. #J-188...
Senior
EITACIES Inc.
Santa Clara, CA
1 day ago
Senior Director, AI-Driven Site Reliability Engineering
JPMorgan Chase & Co. is seeking a Director of Site Reliability Engineering to partner with the Infrastructure Platforms and Foundational Services team in Palo Alto. This role involves guiding stakeholders through complex projects, leading the application of AI capabilities...
Senior
JPMorgan Chase & Co.
Palo Alto, CA
2 days ago
Director, Site Reliability Engineering Sunnyvale, CA , USA
$250k
...systems, eGain provides the single source of truth—explainable, reliable, and maintainable—that serves as the repository for all... ...at scale. Position Overview As Director of Site Reliability Engineering, you will ensure that eGain’s AI knowledge management platform...
Work at office
eGain Corporation
Sunnyvale, CA
4 days ago
Senior Site Reliability Engineer: Cloud, Kubernetes & CI/CD
A leading tech recruiting firm is seeking a Site Reliability Engineer to manage and optimize cloud infrastructure primarily using GCP or AWS. The role involves maintaining high availability through Kubernetes clusters and improving CI/CD pipelines with Terraform. Ideal...
Senior
Amiri Recruiting
Mountain View, CA
1 day ago
Senior Site Reliability Engineer - Cloud AI Infrastructure
Cerebras is looking for a Senior Site Reliability Engineer to join their Infrastructure team in Palo Alto, California. This role involves designing and optimizing infrastructure for distributed AI applications, contributing to the open-source Ray project, and ensuring...
Senior
Cerebras
Palo Alto, CA
2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!