Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Manager, Site Reliability Engineering

$248k - $396.75k

NVIDIA

For over 25 years, NVIDIA has been at the forefront of transforming computer graphics, PC gaming, and accelerated computing, driven by a legacy of continuous innovation and exceptional talent. We are now leveraging the immense potential of AI to usher in the next era of computing, where our GPUs power the "brains" of computers, robots, and autonomous vehicles that can comprehend the world. This pioneering work demands vision, innovation, and the world's best talent. Join our diverse and supportive environment, where NVIDIANs are inspired to excel and make a profound global impact.NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered systems that enhance reliability, speed, and employee experience. We offer an outstanding opportunity to lead and refine Incident, Problem, and Change Management into an intelligent, automated operating model using observability, AI insights, and orchestration. This leader will apply strong operational execution with an SRE attitude, facilitating the move from reactive processes to predictive and autonomous operations.What you’ll be doingManage the full lifecycle of Incident, Problem, and CM as a 247 operational function, ensuring high reliability and minimal business disruption.Transform incident response by bringing to bear AI detection, correlation, and guided remediation, reducing time to detect, respond, and resolve.Build and scale intelligent incident workflows that integrate monitoring, telemetry, and service context to enable faster and more consistent response.Evolve Problem Management into a data-driven field, using AI and analytics to identify patterns, eliminate recurring issues, and drive systemic fixes.Modernize CM by introducing risk-aware, data-driven decisioning, improving change success rates, and reducing blast radius.Drive the adoption of observability as a foundation, ensuring service-level visibility, signal quality, and actionable insights across the IT ecosystem.Lead the development of automation and orchestration platforms that reduce manual effort across the outage lifecycle, including detection, triage, communication, and RCA or equivalent experience.Partner closely with engineering, infrastructure, and business teams to align operations with service reliability goals and SLOs.What we need to see:BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience).5+ years of experience leading and managing global IT operations or service management teams, with growing scope and complexity.12+ overall years of experience in Site Reliability Engineering, IT Service Management, with a focus on Incident Management, Problem Management, and Configuration ManagementProven proficiency in Incident, Problem, and CM with a consistent record of delivering measurable gains in reliability and efficiency.Demonstrated experience applying AI, automation, or advanced analytics to improve operational outcomes.Solid understanding of observability, monitoring ecosystems, and modern reliability practices (SRE principles, SLOs, error budgets).Demonstrated ability to move organizations from process-heavy to technology-focused operating models.Strong leadership capability with experience building and scaling engineering-focused teams (SRE, SWE, or equivalent).Ability to deliver executive-level communication and insights, translating operational signals into clear, actionable narratives for leadership.Ability to build and lead a high-performing team of SREs and engineers, encouraging a culture of ownership, innovation, and continuous improvement.Ways to stand out from the crowd:ITIL knowledge and/or certificationExperience building or scaling AI-powered operational platforms.Ability to challenge traditional ITSM models and introduce innovative, scalable approaches.A mentality passionate about automation first, prevention over reaction, and systems over process.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative, results-oriented and enjoy learning while having fun, then what are you waiting for? Apply today!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 248,000 USD - 396,750 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Senior Manager, Site Reliability Engineering in Santa Clara, CA vacancy
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...internal tooling for system deployment, management and maintenance.What You’ll DoOperate and...  ...services, workloads, and platform reliability.You6+ years of experience in a SRE, operations... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ..., simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure...  ...CI/CD: Partner with feature teams to refine Change Management and CI/CD pipelines, ensuring code moves from "commit" to... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    3 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Sunnyvale, CA
    3 days ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external...  ...developing software platforms and frameworks, capacity management and launch reviews. Be part of an on call rotation to... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually...  ..., and refinement.Participate in defining, evolving, and managing SLOsWrite code and automation to reduce operational... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    1 day ago
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...for system deployment, management and maintenance.What You'll...  ...networking teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    12 hours ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...performance, data integrity, and safe change management. You’ll own SLOs/SLIs, incident response... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    12 hours ago
  • $168k - $270.25k

     ...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing...  ....Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    12 hours ago
  • $90k - $180k

     ...colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are seeking a highly skilled and mission-driven Senior... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    12 hours ago
  • $267k - $356k

     ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...the industry, which means reliability and performance aren't just...  ...across new and existing sites using tools such as Ansible,...  ...including integrating with their management and data-plane APIs.Strong... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    12 hours ago
  • $101k - $161k

     ...prestigious awards, such as Best Engineering Team, Best Company for...  ...Work WithWe’re looking for Site Reliability Engineers to join our...  ...CloudVision is an enterprise network management and streaming telemetry...  ...level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    1 day ago
  • $160k - $240k

     ...times a day - quickly, reliably, and securely. Any time...  ...Fiserv.Job TitleSenior Site Reliability EngineerWhat...  ...Site Reliability Engineer do at Fiserv?You will join...  ...define SLIs and SLOs, manage error budgets and translate...  ...or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    2 days ago
  • $146.7k - $339.3k

     ...not available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid...  .... Able to participate in on-call shifts and incident management and work after hours/weekends for application releases/... 
    Senior
    Full time
    Work at office
    Remote work
    Worldwide
    Shift work
    Weekend work

    Zoom

    San Jose, CA
    3 days ago
  • $210.6k - $305.1k

     ...excellence. Drive strategic vision for the management and continued expansion of FedRAMP-...  ...You have led a distributed team of 5+ engineers, can demonstrate strong technical vision...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    12 hours ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    3 days ago
  •  ...across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational...  ...Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes.Have experience with... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $168k - $270.25k

     ...can make a lasting impact on the world. NVIDIA Base Command Manager powers thousands of clusters worldwide, varying from a few...  ...providing excellent, comprehensive support to our customers! ​Sr Site Reliability Engineer in this role will significantly impact and contribute to... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $187.04k - $359.72k

     ..., user interaction, capital management, tax and exchange optimization...  ...for changes that improve reliability and velocity. Qualifications...  ...Computer Science, Electrical Engineering, Computer Engineering or related...  ...Functions and more. On-site presence across teams allows... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    3 days ago
  • Northrop Grumman is seeking a Principal or Senior Principal Technical Services Project Manager to support the Columbia Dreadnought Launcher Program in Sunnyvale, CA. The role requires coordinating with multiple stakeholders and leading project teams in a manufacturing environment... 
    Senior

    Northrop Grumman

    Sunnyvale, CA
    1 day ago
  • $166k - $244k

    Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-...  ...automation. On the SRE team, you\'ll have the opportunity to manage the complex challenges of scale which are unique to Google... 
    Senior
    Full time

    Google

    Sunnyvale, CA
    3 days ago
  • $224k - $356.5k

     ...see how you can make a lasting impact on the world.As a Senior Developer Relations Manager for Data Platforms, you’ll work with our most strategic...  ...offerings and alignment at the business, product, and engineering levels.Build relationships with executive and technical... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $222k - $300.5k

     ...the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-based...  ...negotiable.The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to lead... 
    Senior
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    8 hours ago
  • $224k - $356.5k

     ...accelerated computing and Agentic AI. The Developer Relations Manager is a high-profile role in NVIDIA. In this important role, you...  ...and drive joint GTM with partners.Partner with NVIDIA’s product/engineering teams to influence the NVIDIA product and roadmap development... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...the ecosystem to bring groundbreaking innovation on NVIDIA’s platform.Collaborate with cross-functional teams, including engineering, product management, and NVIDIA’s Govt. relations teams, to drive tasks aligned with our mission.Lead and contribute to industry... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...shared product foundation.What you’ll be doing:Lead the platform engineering team (frontend, backend, data, security) while guiding the...  ...data plane, plugin API) with one accountable owner per question.Manage the external-framing/shareability story: present honestly as... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $308k

     ...a passion for driving new technologies to market, then join us!What you’ll be doing:Lead a team of Product and Product Marketing managers to complete business critical strategiesOwn product definition and positioning of NVIDIA laptop and embedded products with precisionRepresent... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $304k

     ...growth marketing, or demand generation, with at least 5 years in a senior leadership role.Bachelors, and Master’s degree or equivalent...  ....Ways to stand out from the crowd:Background as a software engineer, data scientist, or AI/ML practitioner — you understand how developers... 
    Senior
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally...  ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at...  ...building, deployment, and lifecycle management using frameworks like TensorFlow,... 
    Senior

    JP Morgan Chase

    Palo Alto, CA
    3 days ago
  • $184k - $287.5k

     ...transform industries and improve lives. We’re looking for a Senior Developer Relations Manager to own developer and partner adoption of the NVIDIA...  ...Master’s, or PhD degree in Computer Science, Robotics, Engineering, or a related technical field — or equivalent... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    12 hours ago
  • $184k - $287.5k

     ...systems.We are seeking a highly technical and strategic Senior Developer Relations Manager to lead developer ecosystem engagement for Autonomous...  ...automotive technologies. You will work closely with product, engineering, solutions architecture, sales, marketing, and regional... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Manager, Site Reliability Engineering. Be the first to apply!