Senior Site Reliability Engineer
NVIDIA
We are now looking for a Sr. Site Reliability Engineer (SRE)! NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s motivated by outstanding technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. NVIDIA is at the forefront of generative AI models, from language to images. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work.
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its cloud service team for supporting, triaging, and building Geforcenow cloud gaming platform. As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. We live SRE practices that are key to product quality, such as limiting time spent on reactive operational work, blameless postmortems, proactive identification of potential outages, and iterative improvements, which all make for interesting and dynamic day-to-day work. The person in this position will be responsible for Service Response and workflow and will drive tools/service development to maintain and improve service SLOs. We partner with Service Owners to drive the reliability of the service. Continuously evaluate operational processes, identify opportunities for improvement, and build custom tools, automation, and self-service solutions that improve the overall reliability and operational efficiency of GeForce NOW.
What You Will Be Doing
Monitor, support, and maintain the reliability, availability, and performance of large-scale GeForce NOW production services running across cloud and datacenter environments.
Participate in production incident triage, troubleshooting, and resolution of complex infrastructure and application issues. Take part in the team's on-call rotation, including occasional weekend coverage, to ensure timely restoration of customer-facing services.
Monitor service health using metrics, logs, traces, and dashboards, and proactively identify reliability, performance, and capacity issues before they impact customers.
Collaborate with software engineering, platform, networking, and infrastructure teams to improve operational readiness, reliability, and service resilience.
Drive observability initiatives by improving monitoring, alerting, dashboards, and telemetry to enable faster detection and diagnosis of production issues.
Scale services sustainably by building automation, eliminating operational toil, and continuously improving deployment, recovery, and operational workflows.
Lead and participate in incident response, root cause analysis, and blameless postmortems, driving corrective and preventive actions to improve long-term service reliability.
Design and develop custom tools, automation, and self-service solutions that simplify operations, improve engineer productivity, and enhance the overall GeForce NOW platform.
Continuously evaluate existing operational processes and identify opportunities to improve service reliability, operational efficiency, and customer experience through engineering-driven solutions.
Contribute to the design, deployment, and operation of Kubernetes-based services, ensuring they meet scalability, reliability, and performance requirements.
What we need to see:
BS degree in Computer Science, Computer Engineering, Information Technology, or a related technical field (or equivalent experience).
5+ years of experience supporting and operating mission-critical production services in a live-site environment as a Site Reliability Engineer (SRE), Production Engineer, or similar role.
Strong understanding of containerization, microservices architecture, and Kubernetes, including Kubernetes ecosystem components and operational best practices.
Demonstrated ability to troubleshoot complex production issues, identify root causes, and drive issues to resolution.
Strong understanding of distributed systems and how complex production environments interact across applications, infrastructure, networking, and cloud services.
Experience supporting production operations, including incident management, change management, postmortem reviews, and operational excellence initiatives.
Hands-on experience developing automation using Python, Go, Bash, or similar scripting/programming languages.
Strong understanding of SLOs, SLIs, error budgets, KPIs, and service reliability best practices.
Experience with observability platforms such as Prometheus, Grafana, ELK/OpenSearch, and modern monitoring and alerting solutions.
Experience operating services in public cloud environments such as AWS, Azure, GCP, or equivalent cloud platforms.
Ways to stand out from the crowd:
Experience supporting large-scale customer-facing cloud or gaming services.
Strong Kubernetes operational and troubleshooting expertise.
Experience with observability platforms, including Prometheus, Grafana, ELK/OpenSearch, and OpenTelemetry.
Strong scripting or programming skills in Python, Go, or similar languages with a focus on automation.
Experience driving production incident response, postmortems, and operational excellence initiatives.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you.
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...SeniorFull time
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
- ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will...SeniorRemote work
$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...Senior$65 - $75 per hour
DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role...SeniorRemote work- IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SeniorWork at officeImmediate start
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...SeniorFlexible hours
$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers...SeniorFull time$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
$152.6k - $191.5k
...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include... ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior...SeniorFull timeWork at officeDay shift$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SeniorFull timeWork at officeLocal areaRemote workWork from home$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...SeniorFull time- Job Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering...SeniorFull time
$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind...SeniorFlexible hours- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SeniorTemporary workCasual workWorldwide
$90k - $180k
...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We...SeniorRemote work$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours- ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The... ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and...Senior
$15k
...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorWork at officeLocal areaRemote work$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSenior- ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS...SeniorTemporary work
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SeniorFull timeWork at office2 days per week
$160k - $200k
...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident...SeniorTemporary workWork at officeLocal areaFlexible hours3 days per week- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the...SeniorFull timeWork at office
$104.9k - $174.7k
Are you passionate about improving reliability, scalability, and resilience in complex database... ....Own prioritization of reliability engineering tasks within team backlogs.Lead incident... ...a Service (IaaS).Background in DevOps, site reliability engineering practices, or related...SeniorFull timeLocal area$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability...SeniorOngoing contractLocal area3 days per week- Job Description:Note: Fidelity will not provide immigration sponsorship for this positionThe RoleOur Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high...SeniorFull time
$80k - $140k
Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and...SeniorFull timeFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote United States
- site reliability engineer sre United States
- site reliability engineering manager United States
- site reliability engineer United States
- lead site reliability engineer United States
- senior groundskeeper United States
- senior maintenance supervisor United States
- senior operations associate United States
- senior safety specialist United States
- lcb senior living United States
