Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Manager, Site Reliability Engineering

$227.2k - $324.5k

Tubi TV

About the Role:Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation.We are seeking an experienced and visionary Senior SRE Manager to lead and grow our newly built Site Reliability Engineering team. You are more than a people manager or a tech lead; you are the strategic leader responsible for architecting our reliability roadmap. You will build and mentor a team of talented engineers, foster a culture of blameless learning and continuous improvement, and champion the engineering practices that allow us to balance rapid innovation with rock-solid stability. You will be a key influencer in our engineering leadership, partnering with peers across the organization to ensure reliability is a shared responsibility and a core tenet of our engineering culture.What You'll Do:Team Leadership & Mentorship:Lead, mentor, and grow a team of Site Reliability Engineers. Foster a culture of innovation and technical excellence where engineers feel empowered to do their best work. Provide personalized coaching, create professional development plans, and guide the careers of senior and emerging talent within the team.Establish equitable, sustainable on-call practices (including global coverage where applicable) that protect focus time and avoid burnout.Define team rituals - runbook reviews, game days, and incident retros - that reinforce quality and learning.Strategic Planning & Vision: Define and drive the multi-year technical strategy and vision for Tubi’s observability, and automation platforms. Partner with infra lead to align Tubi’s infrastructure & SRE roadmap. Partner with tech leaders to align the SRE roadmap with business objectives. Champion a data-driven approach to reliability, using Service Level Objectives (SLOs) and error budgets to facilitate productive conversations about risk and feature velocity.Operational Excellence & Incident Management:Own the end-to-end availability, performance, and efficiency of our critical user-facing services. Evolve our incident response practice to reduce Mean Time to Resolution (MTTR) and Mean Time Between Failures (MTBF). Champion a rigorous, blameless, and data-driven post-mortem culture to ensure we learn from both successes and failures, driving eng teams for systemic fixes and automation to prevent the recurrence of incidents.Streamline and improve our existing processes and practices, and collaborate with other teams to enhance our production release standards by improving current processes.Define and tune a 247 on-call rotation for low noise and fast response; act as executive escalation partner during major incidents.Own disaster-recovery strategy (playbooks, failover drills, recovery simulations) and track SLO gaps with time-bound remediations.Financial & Vendor Management: Own the SRE budget, tooling, and headcount. Manage relationships with key third-party vendors for our observability and SRE related AI platforms, work with infra lead and finance team for contract negotiations and ensure we derive maximum value from our investments.Cross-Functional Collaboration: Act as a key influencer and strategic partner to leaders in Software Engineering, Product Management, and Infra/Sec. Drive the adoption of SRE best practices and principles throughout the organization, ensuring new services are designed for reliability, scalability, and observability from day one.The AI Mandate: Building the Future of Observability with AI. You will not just manage a team that uses AI; you will lead the charge in building an AI-native SRE function. This is a strategic mandate that requires a forward-thinking leader who understands both the potential and the pitfalls of integrating intelligent systems into critical operations. This includes:AIOps Strategy Development: Developing and executing the strategy for integrating AIOps and machine learning into our observability stack. Your goal will be to move the team from a reactive monitoring posture to one of predictive maintenance and automated anomaly detection, fundamentally changing how we ensure reliability.Accelerating Automation with AI: Championing the effective and responsible use of AI-assisted coding tools (e.g., Claude Code, Cursor) within the SRE team. You will set the standards and practices to leverage these tools to accelerate the development of automation, operational tooling, and infrastructure code.Building the Business Case: Building the techno-economic case for new AI tooling, managing vendor relationships, and ensuring the cost-effective and secure implementation of these powerful systems. You must be able to articulate the ROI of these investments in terms of reduced downtime, improved operational efficiency, and faster incident resolution.Fostering Critical AI Literacy: Fostering a culture that can critically evaluate, debug, and learn from the outputs of AI systems. This involves extending our blameless post-mortem philosophy to AI-driven actions and recommendations, ensuring that the team remains in control and understands the "why" behind automated decisions.Your Background:8+ years of experience in a technical field, with at least a year in an engineering leadership position managing SRE, DevOps, or Production Engineering teams.A deep, principled understanding of SRE tenets, including Service Level Indicators (SLIs), SLOs, error budgets, toil reduction, and capacity planning.Exceptional communication, negotiation, and influencing skills, with the ability to articulate complex technical concepts and strategies to both technical and non-technical stakeholders at all levels of the organization.A strong technical background as a hands-on software engineer or site reliability engineer prior to moving into management. Deep knowledge of AWS services (especially networking, IAM, EKS, ALBs/NLBs, Route 53, CloudWatch). Proven experience with Kubernetes in production (EKS preferred), including service exposure, networking, and availability engineering.Hands-on familiarity with modern SRE tools and technologies, including Infrastructure as Code (e.g., Terraform, Ansible), container orchestration (Kubernetes), observability platforms (e.g., Prometheus, Grafana, Datadog, Splunk), and incident tooling (e.g., PagerDuty, FireHydrant), deployment-safety tooling (e.g., Argo Rollouts, LaunchDarkly), and observability standards (e.g., OpenTelemetry).#LI-BT1#LI-Hybrid Pursuant to state and local pay disclosure requirements, the pay range for this role, with final offer amount dependent on education, skills, experience, and location, is listed annually below. This role is also eligible for various benefits, including medical/dental/vision, insurance, a 401(k) plan, paid time off, and other benefits in accordance with applicable plan documents.High cost labor markets such as but not limited to Los Angeles, New York City, and San Francisco$227,200—$324,500 USDTubi is a division of Fox Corporation, and the FOX Employee Benefits summarized here, covers the majority of all US employee benefits. The following distinctions below outline the differences between the Tubi and FOX benefits:For US-based non-exempt Tubi employees, the FOX Employee Benefits summary accurately captures the Vacation and Sick Time.For all salaried/exempt employees, in lieu of the FOX Vacation policy, Tubi offers a Flexible Time off Policy to manage all personal matters.For all full-time, regular employees, in lieu of FOX Paid Parental Leave, Tubi offers a generous Parental Leave Program, which allows parents twelve (12) weeks of paid bonding leave within the first year of birth, adoption, surrogacy, or foster placement of a child in addition to applicable government leave program(s) and FOX’s short-term disability policy. This time is 100% paid through a combination of any applicable state, city, and federal leaves and wage-replacement programs in addition to contributions made by Tubi.For all full-time, regular employees, Tubi offers a monthly wellness reimbursement.About Tubi:Boldly built for every fandom, Tubi is a free streaming service that entertains over 100 million monthly active users. Tubi offers the world's largest collection of Hollywood movies and TV shows, thousands of creator-led stories and hundreds of Tubi Originals made for the most passionate fans. Headquartered in San Francisco and founded in 2014, Tubi is part of Tubi Media Group, a division of Fox Corporation.We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, gender identity, disability, protected veteran status, or any other characteristic protected by law. We will consider for employment qualified applicants with criminal histories consistent with applicable law.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Manager, Site Reliability Engineering in San Francisco, CA vacancy
  • $165k - $225.6k

     ...to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. THE SENIOR SITE RELIABILITY ENGINEER OPPORTUNITY Reporting to the Manager, Site Reliability Engineering, this role will help build, improve, and... 
    Senior
    Permanent employment
    Full time
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and...  ...cross-functionallyNice-to-HaveExperience with cloud and managed services (e.g. AWS)Experience supporting data-intensive platforms... 
    Senior

    Alembic

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and...  ...Terraform experience, including authoring reusable modules, managing state and environments, and delivering infrastructure... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...practices such as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction.... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior...  ...and performanceImplement and manage service-level indicators (SLIs), objectives... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    San Francisco, CA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range...  ...observability and alerting systems.The Fleet Management team provides the core runtime...  ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  • $159.2k - $301.6k

     ...APIs for two primary uses: (1) managing Graph plugins, their...  ...Graphs on the cloud. In this reliability-focused role, you will own the...  ...'ll partner with the backend engineers building these APIs to make sure...  ...5-10 years of experience in site reliability engineering, infrastructure... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    3 days ago
  •  ...fully integrated solutions to manage everything from business...  ...what’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...together to build scalable, reliable, and secure products that...  ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...  ...toil and improve operational efficiency.Incident Management: Lead the coordinated response to incidents as an... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  • $165k - $241.4k

     ...performance, efficiency, change management, monitoring, emergency...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    1 day ago
  • $220k - $235k

     ...Quadrant for Contract Lifecycle Management, a Fortune Great Place to...  ..., high-output Staff/Senior Staff SRE to define the future...  ...cloud platform and champion engineering excellence across Ironclad....  ...strategic direction for the Site Reliability Engineering team and our... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    1 day ago
  •  ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on....  ...workloads.Own network device configuration management end to end, ensuring consistency and... 
    Senior

    Alembic

    San Francisco, CA
    4 days ago
  • $215k - $275k

     ...raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide...  ...for both the control plane, which orchestrates cluster management, scheduling, and user access, and the data plane, which... 
    Senior
    Work at office

    Anyscale

    San Francisco, CA
    4 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure...  ..., GCP), including network and compute security, identity management, and cloud security posture management (CSPM)Automation and... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    2 days ago
  • $165k - $241.4k

     ...and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud,... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $210.8k - $272.8k

     ...millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on...  ...for deployment, change, service, and infrastructure management Troubleshoot and debug critical systems throughout the SDLC... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    2 days ago
  • $175k - $250k

     ...250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA...  ...unavailable. Modality: On-Site only. Must live within...  ...scalability, performance, and reliability across environments. What...  ...AI workloads at scale Manage and automate GPU compute clusters... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    3 days ago
  • $266k - $398k

     ...Director, Site Reliability Engineering – Infrastructure Platform Okta is The World’s Identity Company. Okta provides secure access, authentication...  ...mentor, and grow a high‑performing team of engineers and managers across platform, infrastructure, and shared services domains... 
    Senior
    Permanent employment
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    3 days ago
  • $150k - $184k

    Senior Manager, Talent Management San Francisco, CA At Rothy’s, we know there’s a better way to do business, and it starts by putting the planet and its people first. More than 225 million single-use plastic bottles and 715,000 pounds of ocean-bound marine plastic have... 
    Senior
    Full time
    Flexible hours

    Rothy's

    San Francisco, CA
    8 hours ago
  •  ...Job Description Role Summary The Senior Project Scheduler leads the development,...  ...contractor submissions. EVM & Change Management: Apply Earned Value Management (EVM) practices...  ...schedule comparison reports and updates for site and executive leadership.... 
    Senior
    For contractors
    Shift work

    Scanr Solutions

    San Francisco, CA
    2 days ago
  • $245k - $295k

     ...part of a high-performing team that believes in each other, come build with us at Crusoe.About the Role:Join Crusoe as a Senior Engineering Manager and lead a talented team focused on revolutionizing our cloud infrastructure. In this pivotal role, you'll lead the Command... 
    Senior
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $124k - $280k

     ...Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design...  ...data storage solutions using cloud services- Designing and managing data warehouses and data lakes- Implementing IAM roles and... 
    Senior
    Full time
    H1b

    PwC

    San Francisco, CA
    1 day ago
  • $15k

     ...worked at the frontier of applying AI/ML to investment management. We have become a multibillion-dollar asset manager,...  ...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    11 hours ago
  • $150k - $185k

     ...global boutique consultancy dedicated to managing and representing our clients’ best...  ...Management, Civil/Structural/Mechanical Engineering, Architecture, Business, or equivalent industry...  ...thorough understanding of MEP systems, site logistics, and building design/... 
    Senior
    Full time
    For contractors
    Local area

    MGAC Private/unlisted

    San Francisco, CA
    8 hours ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help...  ...in advancing the reliability, scalability, automation, observability...  ...is a highly hands-on engineer with strong production experience...  ...provisioning and change management through Terraform/TerragruntBuilding... 
    Senior
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    11 hours ago
  •  ...Read more: careers.bms.com/working-with-us. Position Summary The Senior Manager of Biostatistics is a member of cross-functional development...  .... Eligibility for specific benefits listed on our careers site may vary based on the job and location. For more on benefits,... 
    Senior
    Hourly pay
    Full time
    Temporary work
    Part time
    For contractors
    Summer work
    Live in
    Work at office
    Local area
    Remote work
    Flexible hours
    Shift work

    Bristol Myers Squibb

    Brisbane, CA
    8 hours ago
  • $175k - $190k

     ...A leading insurance brokerage firm is seeking an Experienced P&C Client Executive in San Francisco to manage client relationships and provide expert guidance on insurance programs. Candidates should have a college degree and at least 10 years of experience in a similar... 
    Senior

    EPIC Insurance Brokers & Consultants

    San Francisco, CA
    2 days ago
  •  ...Neura Market is looking for a Senior Manager of Compensation in San Francisco to partner with GTM and G&A teams. This hands-on role involves guiding compensation decisions, using data to drive insights, and working on executive compensation strategies. Ideal candidates... 
    Senior

    Neura Market

    San Francisco, CA
    2 days ago
  •  ...jobs.frontdoordefense.com - Jobboard in San Francisco is hiring a Manager, People Partners to support hardware teams by driving talent strategies. This role entails advising senior leaders, developing hiring strategies, and managing various talent programs. The ideal... 
    Senior
    Work at office
    3 days per week

    jobs.frontdoordefense.com - Jobboard

    San Francisco, CA
    2 days ago
  •  ...seasoned customer success leader to join our Implementation team in San Francisco. In this role, you'll coach a team of Implementation Managers, engage with high priority customers, and drive improvements on key metrics. The ideal candidate will have 5+ years of SaaS... 
    Senior

    SupportFinity

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Manager, Site Reliability Engineering. Be the first to apply!