Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager of Site Reliability Engineering

Okta

Requirements 3+ years of experience in technical leadership & people management Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines Strong background and hands‑on experience in SW development, PaaS and automation Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment Effective verbal, written communication and interpersonal skills Computer Science Degree or related degree or equivalent experience This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire What the job involves As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management Manage service and business expectations and prioritize resource allocation Maintain a deep knowledge of industry best practices, evolving trends, and technologies #J-18808-Ljbffr Okta

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Manager of Site Reliability Engineering in San Francisco, CA vacancy
  •  ...of-the-art AI. As a Senior SRE, you'll tackle the scaling and reliability challenges that come with adding terabytes of data monthly and...  ...Build out distributed tracing, metrics, and alerting that give engineers clear visibility into system behavior and accelerate debugging... 
    Suggested

    Unify

    San Francisco, CA
    1 day ago
  •  ...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging... 
    Suggested
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    1 day ago
  •  ...Open Source LLM Gateway Engineer LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA...  ...expanding and seeking our 6th Engineer focused on owning reliability, performance, and infrastructure stability for the LiteLLM proxy... 
    Suggested

    BerriAI

    San Francisco, CA
    7 hours ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ..., GPU utilization, concurrency, and model lifecycle management. Define and instrument SLOs and SLIs across customer workloads... 
    Suggested
    Flexible hours

    Baseten

    San Francisco, CA
    4 days ago
  • $127k - $249k

    THE TEAM Platform Engineering is the department within SRE that is responsible for a range...  ...and alerting systems. The Fleet Management team provides the core runtime environment...  ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Suggested
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    7 hours ago
  •  ...$10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. About the Role As a Site Reliability Engineer (SRE) at Mercor, you'll own production reliability across our most critical systems, partnering directly with infrastructure... 
    Work at office
    Relocation package

    Mercor Alabaster

    San Francisco, CA
    4 days ago
  •  ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure...  ...) ~ Strong debugging, problem-solving, and incident-management skills Preferred Experience with... 

    Blaxel, Inc

    San Francisco, CA
    2 days ago
  • $350k

     ...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity through advancing collaborative general...  ...have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full... 
    Local area
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    7 hours ago
  •  ...DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you...  ...Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production environments at scale. - Data Proficiency:... 

    BayOne Solutions

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure...  ...such as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction. Success... 
    For contractors

    Autodesk

    San Francisco, CA
    5 days ago
  • $195k - $240k

     ...Senior Site Reliability Engineer San Francisco (Hybrid) At You.com, we are building the AI Search Infrastructure that powers modern AI systems...  ...or non-actionable alerts on a regular cadence. Help manage incident management processes and playbooks. Qualifications... 
    Full time
    Immediate start
    Remote work
    Work from home
    Flexible hours

    Y.O.U.

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...efforts. Job Category Software Engineering Job Details About Salesforce...  ...senior engineering candidate to join the Site Reliability organization in San Francisco. Working...  ...operational efficiency. * Incident Management: Lead the coordinated response to incidents... 
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  •  ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability...  ..., deployment automation, rollback mechanisms, and config management Implement and maintain monitoring, alerting, and... 

    Alembic Technologies

    San Francisco, CA
    4 days ago
  • $181.69k - $213.75k

     ...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects...  ...funds and SPVs, representing nearly $185B in assets under management, with tools designed to enhance the strategic impact of... 
    Full time
    Work at office

    Carta

    San Francisco, CA
    4 days ago
  • $125k - $165k

     ...Site Reliability Engineer TELCOR Inc, a leading innovator in laboratory software, is looking for a Site Reliability Engineer to join our TELCOR...  ...across cloud and containerized environments, as well as manage production infrastructure and deployment workflows across environments... 
    Work at office
    Remote work

    TELCOR

    San Francisco, CA
    1 day ago
  •  ...culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of...  ...6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    1 day ago
  •  ...products that empower people across the globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability,... 
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    1 day ago
  •  ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the...  ...infrastructure for our users that scales, is reliable, and makes the complexities of operating...  ...of the challenges: streaming, token management, rate limits, model-specific quirks.... 
    Permanent employment
    Shift work

    Arena AI

    San Francisco, CA
    2 days ago
  •  ...Superhuman Docs's collaborative workspaces, Mail's inbox management, and Go, the proactive AI assistant that...  ...responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    3 days ago
  • $260k - $300k

     ...makers of Devin, the first AI software engineer, and Windsurf, the AI-native IDE. Together...  .... You will own both the production reliability of our user-facing products and the platform...  ...matters. Infrastructure as Code: Manage cloud infrastructure through code. Build... 

    Cognition AI

    San Francisco, CA
    4 days ago
  •  ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI...  ...Systems Builder — Close the Loop Build and maintain fleet management systems: OTA update pipelines, device health tracking,... 
    Remote work

    Specter Services LLC

    San Francisco, CA
    1 day ago
  • $166.9k - $225.9k

     ...SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a...  ...infrastructure requests: ECS task management, secret rotations, Terraform changes...  ...: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering... 
    Work at office
    Immediate start
    Worldwide
    Monday to Friday
    Flexible hours

    Drata Inc

    San Francisco, CA
    1 day ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations...  ...Key Responsibilities Capacity Ingestion and Management: Takes proactive steps to design and architect infrastructure... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    7 hours ago
  •  ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally... 
    Contract work
    Local area

    InterSources

    San Francisco, CA
    4 days ago
  • $230k - $310k

     ...daily users while enabling our engineering teams to ship fast. You'll...  ...automation and tooling that improves reliability and partnering with...  ...scale with the product Manage and optimize our compute, networking...  ...'ll Bring ~5+ years in site reliability engineering,... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    3 days ago
  • $159.2k - $301.6k

     ...APIs for two primary uses: (1) managing Graph plugins, their...  ...Graphs on the cloud. In this reliability-focused role, you will own the...  ...'ll partner with the backend engineers building these APIs to make sure...  ...~5-10 years of experience in site reliability engineering, infrastructure... 
    Temporary work
    Local area
    Worldwide

    Adobe

    San Francisco, CA
    4 days ago
  • $287k

     ...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to...  ...hit our SLAs. What? We're looking for a Senior or Staff Site level Reliability Engineer as part of Infrastructure team to: Own uptime... 
    Contract work
    Work at office
    Remote work

    IVO Inc

    San Francisco, CA
    4 days ago
  • $220k - $235k

     ...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts...  ...Wave and Gartner Magic Quadrant for Contract Lifecycle Management, a Fortune Great Place to Work, and one of Fast Company's... 
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    1 day ago
  • $200k - $260k

     ...Infrastructure Team as a technical leader driving reliability, automation, and scalability across the...  ...practices across teams, mentor senior engineers, and be a primary escalation point for...  ...engineers, without needing formal management authority to do it ~ Strong bias for... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager of Site Reliability Engineering. Be the first to apply!