Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Manager, Site Reliability Engineering

$222k - $300.5k

Intuit Financial Services

Company OverviewIntuit is the global financial technology platform that powers prosperity for the people and communities we serve. With tens of millions of customers worldwide using products such as TurboTax, Credit Karma, QuickBooks, and Mailchimp, we believe that everyone should have the opportunity to prosper. We never stop working to find new, innovative ways to make that possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps QuickBooks, TurboTax, Credit Karma, and Mailchimp running for hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency tooling, and incident response capability that underpins Intuit's money-movement and fintech services — where availability, data integrity, and trust are non-negotiable.The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10–15 systems and reliability engineers responsible for the availability, performance, and operational health of Fintech Platform services running in AWS. This leader owns the strategy and execution behind operational excellence: driving toward a 99.999% availability bar, maturing incident management practices, and building self-healing, well-instrumented infrastructure at scale.This is a player-coach role. You will set technical direction and organizational strategy while staying close to the systems — reviewing designs, joining incident bridges, and coaching engineers through complex production issues. You'll partner closely with software engineering, product, security, and other SRE/infrastructure leaders across Intuit to raise the bar on reliability company-wide.A defining priority for this role is AI Ops: embedding AI-driven, autonomous operations into how the team runs infrastructure. You will lead the shift from manual, human-triggered response toward self-healing systems that detect, diagnose, and remediate issues autonomously — reducing developer toil, cutting MTTR, and freeing engineering capacity to focus on higher-value work. Done well, this delivers 3x the operational impact of the team today and directly accelerates the pace at which we deliver value to customers.ResponsibilitiesResponsibilitiesOwn end-to-end operational excellence for Fintech Platform services: define and drive the strategy for achieving and sustaining 99.999% availability across customer-facing and internal systems.Lead, grow, and directly manage a team of 10–15 systems/site reliability engineers — hiring, mentoring, setting goals, and developing the next generation of technical leaders.Act as a hands-on technical leader: participate in architecture and design reviews, write and review code/IaC where needed, and dive into production systems alongside the team.AI Ops: Driving 3x Impact Through Autonomous Operations- Define and execute an AI Ops roadmap that embeds autonomous detection, diagnosis, and remediation into production systems, targeting a 3x improvementIdentify high-toil, repetitive operational workflows and systematically replace them with autonomous agents and automation, freeing engineers to focus on higher-leverage engineering work.Measure and report on toil reduction, automation coverage, and velocity gains, tying AI Ops investment directly to faster, safer delivery of customer value.Drive incident management maturity — own the incident command process, lead or oversee response for high-severity (P1/P2) incidents, and ensure rigorous root-cause analysis and blameless postmortems.Build and scale AWS cloud infrastructure (compute, networking, storage, container orchestration) with a focus on resiliency, auto-remediation, chaos engineering, and multi-AZ/multi-region failover.Define and report on SLOs/SLIs, error budgets, and availability metrics; use data to prioritize reliability investments and reduce toil through automation.Partner with software engineering, product management, security, and compliance teams to embed reliability, observability, and operational readiness into the software development lifecycle.Establish and continuously improve on-call practices, runbooks, alerting, and escalation paths to reduce MTTD/MTTR.Own capacity planning, cost optimization, and infrastructure roadmap decisions for the systems under your purview.Represent Infrastructure & SRE in leadership forums, change advisory boards, and executive incident reviews; communicate risk and operational posture clearly to senior stakeholders.Champion a culture of operational rigor, psychological safety, and continuous improvement across the team.QualificationsQualifications8+ years of experience in systems engineering, site reliability engineering, or infrastructure engineering, with 3+ years directly managing engineering teams.Proven, hands-on experience operating production infrastructure in AWS at scale (EC2, EKS/ECS, VPC, RDS/DynamoDB, IAM, CloudWatch, Auto Scaling, and related services).Track record of driving high-availability outcomes (99.9%+ and above) for mission-critical, customer-facing systems, ideally in fintech, payments, or another regulated/high-trust domain.Deep experience with incident management — running incident command, leading postmortems, and building organizational muscle around detection, response, and prevention.Strong technical foundation in distributed systems, networking, containerization/orchestration (Kubernetes), and infrastructure-as-code (Terraform, CloudFormation, or similar).Experience with observability and reliability tooling (e.g., Datadog, Splunk, PagerDuty, Prometheus/Grafana) and building SLO-driven operations.Demonstrated ability to balance hands-on technical depth with people leadership — comfortable reviewing a design doc and coaching a direct report in the same day.Excellent communication skills, with experience presenting risk, status, and strategy to senior/executive leadership.Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.Experience designing or scaling AI Ops / autonomous remediation capabilities (AIOps platforms, ML-based anomaly detection, agentic automation) that measurably reduced toil or MTTR.Experience operating within a regulated fintech, banking, or payments environment (PCI, SOC 2, money movement/ACH systems).Prior experience building or scaling a chaos engineering or resilience testing practice.Familiarity with cost and capacity management for large-scale multi-account AWS environments.Experience leading through major incidents involving cross-functional executive stakeholders.Intuit provides a competitive compensation package with a strong pay for performance rewards approach. This position may be eligible for a cash bonus, equity rewards and benefits, in accordance with our applicable plans and programs (see more about our compensation and benefits at Intuit: Careers | Benefits). Pay offered is based on factors such as job-related knowledge, skills, experience, and work location. To drive ongoing fair pay for employees, Intuit conducts regular comparisons across categories of ethnicity and gender. The expected base pay range for this position is: Mountain View $222,000 - $300,500

Vacancy posted 14 days ago
Similar jobs that could be interesting for youBased on the Senior Manager, Site Reliability Engineering in Mountain View, CA vacancy
  • $165k - $280k

     ...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we’re leveraging our experience in...  ...infrastructure to support a multi-region environment Manage petabyte scale bare metal compute clusters Closely collaborate... 
    Senior
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    InvestedintheMission

    Palo Alto, CA
    4 days ago
  •  ...role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining...  ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at...  ...building, deployment, and lifecycle management using frameworks like TensorFlow, PyTorch... 
    Senior

    J.P. Morgan

    Palo Alto, CA
    6 hours ago
  • $160k - $240k

     ...times a day - quickly, reliably, and securely. Any time...  ...Fiserv.Job TitleSenior Site Reliability EngineerWhat...  ...Site Reliability Engineer do at Fiserv?You will join...  ...define SLIs and SLOs, manage error budgets and translate...  ...or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    18 days ago
  • $262k - $364k

     ...within the AViD ecosystem have reliability and uptime appropriate to...  ...and performance.Build creative engineering solutions to operations and infrastructure...  ...in a strategic way.Site Reliability Engineering (SRE)...  ...’ll have the opportunity to manage the complex challenges of... 
    Senior

    Google

    Mountain View, CA
    7 days ago
  • $262k - $364k

     ...a team of Software/Systems Engineers on projects for users and be...  ...quality technical execution.Manage on-call rotations across continents...  ..., or a related field.Site Reliability Engineering (SRE) combines software...  ...chose to join SRE.As the Senior Engineering Manager for... 
    Senior

    Google

    Sunnyvale, CA
    14 days ago
  •  ...EarnIn in Mountain View is seeking an experienced full-stack engineering leader to head the Early Bets team. You will balance hands-on coding with team management, guiding architecture and delivery of early-stage products designed to boost retention and revenue. You... 
    Senior

    Jobleads-US

    Mountain View, CA
    4 days ago
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 
    Senior

    Oracle

    Santa Clara, CA
    3 days ago
  •  ...transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud...  ...Feed the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated... 
    Senior
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    2 days ago
  •  ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable...  ...role combines database engineering, site reliability engineering, Linux...  ...production and cloud environments. Manage PostgreSQL deployments on Kubernetes... 
    Senior

    Neshent Technologies

    Los Gatos, CA
    2 days ago
  •  ...(eCF) Business Data Technologies group seeks a Sr Tech Program Manager in Sunnyvale to lead cross-organizational programs that scale and...  ...into actionable programs, coordinating with tech leadership, engineers, and stakeholders across stores, devices and other teams. You... 
    Senior

    Amazon

    Sunnyvale, CA
    4 days ago
  • $180k - $230k

     ...Description We're looking for a Senior SRE to own the reliability, scalability, and observability of...  ...work closely with platform and data engineering to keep high-throughput, data-intensive...  ...— deployment pipelines, capacity management, self-healing systems Partner... 
    Senior
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    1 day ago
  • $145k - $175k

     ...Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability,...  ...and review architectures for operational risk. Manage Kubernetes or container orchestration environments at scale... 
    Senior
    Full time
    Remote work

    GrabJobs

    Santa Clara, CA
    6 hours ago
  • $187.04k - $359.72k

     ..., user interaction, capital management, tax and exchange optimization...  ...for changes that improve reliability and velocity. Qualifications...  ...Computer Science, Electrical Engineering, Computer Engineering or related...  ...Functions and more. On-site presence across teams allows... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    3 days ago
  • $130k - $180k

     ...You will report to the Senior Director of Sales,...  ...sales organization, sales engineering, as well as...  ...customers Account Management: Once a program is underway...  ...financial support ~ Off-sites and many social events...  ...institutions with a fast, reliable, and flexible way to... 
    Senior
    Full time
    Contract work
    Temporary work
    Work at office
    Remote work
    Relocation package
    Flexible hours

    Loft Orbital Solutions

    Sunnyvale, CA
    1 day ago
  • DPR Construction is seeking a seasoned Project Controls Manager in Palo Alto, CA to lead planning, scheduling, cost control, and change management across multiple projects. You will implement project controls methodologies, develop budgets and forecasts, monitor performance... 
    Senior

    DPR Construction

    Palo Alto, CA
    5 days ago
  •  ...Senior Electrical Project ManagerOur client, a growing electrical construction contractor, is seeking a Senior Project Manager who thrives in a hands-on, operational role. This position oversees all aspects of project execution for commercial and industrial electrical... 
    Senior
    For contractors
    For subcontractor

    gpac

    Sunnyvale, CA
    2 days ago
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...for system deployment, management and maintenance.What You'll...  ...networking teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  •  ..., simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure...  ...CI/CD: Partner with feature teams to refine Change Management and CI/CD pipelines, ensuring code moves from "commit" to... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    a month ago
  • $150k - $185k

    An innovative tech company in Palo Alto is seeking a Product Marketing Manager to shape the company's messaging, competitive positioning, and sales enablement. This foundational role will require independent thinking and the ability to thrive in a startup environment. Ideal... 
    Senior

    Jobleads-US

    Palo Alto, CA
    6 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...internal tooling for system deployment, management and maintenance.What You’ll DoOperate and...  ...services, workloads, and platform reliability.You6+ years of experience in a SRE, operations... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    22 days ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...performance, data integrity, and safe change management. You’ll own SLOs/SLIs, incident response... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $168k - $270.25k

     ...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing...  ....Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    25 days ago
  • $300k - $350k

     ...Senior Account Executive DeepInfra is building the foundation...  ...modern AI in production — simply, reliably, and at scale. Our team has...  .... Proactively build and manage pipeline by prospecting into...  ...Work closely with GTM and Engineering to shape how we sell and support... 
    Senior
    Contract work
    Shift work

    Deepinfra

    Palo Alto, CA
    2 days ago
  • $267k - $356k

     ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...the industry, which means reliability and performance aren't just...  ...across new and existing sites using tools such as Ansible,...  ...including integrating with their management and data-plane APIs.Strong... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    25 days ago
  •  ...Mechanical Senior Project ManagerMechanical Senior Project Manager needed ASAP for a well-respected, growing mechanical contractor in California! Come work for a company that has many years of proven excellence, great compensation and benefits that shows no signs of slowing... 
    Senior
    For contractors
    Immediate start

    gpac

    Sunnyvale, CA
    4 days ago
  • $61k - $101k

     ...Requirements: We require formal training or certification in site reliability engineering, along with 5+ years of hands-on experience. We need...  ...with AI/ML model building, deployment, and lifecycle management using frameworks like TensorFlow, PyTorch, or scikit-... 
    Senior
    Full time

    J.P. Morgan

    Palo Alto, CA
    17 days ago
  • Next Frontier Capital is seeking a Relationship Manager for JPMorgan Private Client to foster high-net-worth client relationships in Palo Alto, California. The successful candidate will have over 10 years in Financial Services and demonstrate an entrepreneurial spirit.... 
    Senior

    Next Frontier Capital

    Palo Alto, CA
    4 days ago
  •  ...Job Description Job Description Senior Super Product Design Manager Super Company is seeking a Senior Super Product Design Manager to lead high-impact design efforts and help shape exceptional product experiences. Position Summary The Senior Super Product Design... 
    Senior
    Full time
    Work at office

    Super Company

    Mountain View, CA
    26 days ago
  •  ...Wealthfront is seeking an experienced Senior Marketing Manager to elevate our paid partnerships strategy and accelerate customer acquisition...  ...paid marketing team, and partners across the company in Product Marketing, Creative, Comms, DS and Engineering, among others.... 
    Senior
    Full time

    Wealthfront

    Palo Alto, CA
    11 days ago
  • $262k - $364k

    Architect and scale resilient runtime engines capable of state persistence, long-context working...  ...zero-shot tool selection, OAuth/identity managing, and API integration for short-running...  ...of large-scale projects across multiple sites internationally.The Search Verticals team... 
    Senior

    Google

    Mountain View, CA
    19 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Manager, Site Reliability Engineering. Be the first to apply!