Manager of Site Reliability Engineering
Okta
Requirements 3+ years of experience in technical leadership & people management Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines Strong background and hands‑on experience in SW development, PaaS and automation Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment Effective verbal, written communication and interpersonal skills Computer Science Degree or related degree or equivalent experience This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire What the job involves As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management Manage service and business expectations and prioritize resource allocation Maintain a deep knowledge of industry best practices, evolving trends, and technologies #J-18808-Ljbffr Okta
- ...of-the-art AI. As a Senior SRE, you'll tackle the scaling and reliability challenges that come with adding terabytes of data monthly and... ...Build out distributed tracing, metrics, and alerting that give engineers clear visibility into system behavior and accelerate debugging...Suggested
- ...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging...SuggestedWorldwideShift work
- ...Open Source LLM Gateway Engineer LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA... ...expanding and seeking our 6th Engineer focused on owning reliability, performance, and infrastructure stability for the LiteLLM proxy...Suggested
- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ..., GPU utilization, concurrency, and model lifecycle management. Define and instrument SLOs and SLIs across customer workloads...SuggestedFlexible hours
$127k - $249k
THE TEAM Platform Engineering is the department within SRE that is responsible for a range... ...and alerting systems. The Fleet Management team provides the core runtime environment... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...SuggestedWork at officeLocal areaRemote workWorldwideFlexible hours- ...$10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. About the Role As a Site Reliability Engineer (SRE) at Mercor, you'll own production reliability across our most critical systems, partnering directly with infrastructure...Work at officeRelocation package
- ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure... ...) ~ Strong debugging, problem-solving, and incident-management skills Preferred Experience with...
$350k
...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity through advancing collaborative general... ...have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full...Local areaVisa sponsorshipWork visaRelocation package- ...DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you... ...Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production environments at scale. - Data Proficiency:...
$117k - $209.33k
...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure... ...such as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction. Success...For contractors$195k - $240k
...Senior Site Reliability Engineer San Francisco (Hybrid) At You.com, we are building the AI Search Infrastructure that powers modern AI systems... ...or non-actionable alerts on a regular cadence. Help manage incident management processes and playbooks. Qualifications...Full timeImmediate startRemote workWork from homeFlexible hours$148.5k - $223.9k
...efforts. Job Category Software Engineering Job Details About Salesforce... ...senior engineering candidate to join the Site Reliability organization in San Francisco. Working... ...operational efficiency. * Incident Management: Lead the coordinated response to incidents...WorldwideWeekend work- ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability... ..., deployment automation, rollback mechanisms, and config management Implement and maintain monitoring, alerting, and...
$181.69k - $213.75k
...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects... ...funds and SPVs, representing nearly $185B in assets under management, with tools designed to enhance the strategic impact of...Full timeWork at office$125k - $165k
...Site Reliability Engineer TELCOR Inc, a leading innovator in laboratory software, is looking for a Site Reliability Engineer to join our TELCOR... ...across cloud and containerized environments, as well as manage production infrastructure and deployment workflows across environments...Work at officeRemote work- ...culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of... ...6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-...Immediate startRemote workWorldwide
- ...products that empower people across the globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability,...Temporary workWorldwide
- ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the... ...infrastructure for our users that scales, is reliable, and makes the complexities of operating... ...of the challenges: streaming, token management, rate limits, model-specific quirks....Permanent employmentShift work
- ...Superhuman Docs's collaborative workspaces, Mail's inbox management, and Go, the proactive AI assistant that... ...responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth...WorldwideHome officeFlexible hours
$260k - $300k
...makers of Devin, the first AI software engineer, and Windsurf, the AI-native IDE. Together... .... You will own both the production reliability of our user-facing products and the platform... ...matters. Infrastructure as Code: Manage cloud infrastructure through code. Build...- ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI... ...Systems Builder — Close the Loop Build and maintain fleet management systems: OTA update pipelines, device health tracking,...Remote work
$166.9k - $225.9k
...SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a... ...infrastructure requests: ECS task management, secret rotations, Terraform changes... ...: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering...Work at officeImmediate startWorldwideMonday to FridayFlexible hours$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations... ...Key Responsibilities Capacity Ingestion and Management: Takes proactive steps to design and architect infrastructure...Temporary workImmediate startFlexible hoursShift work- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...Contract workLocal area
$230k - $310k
...daily users while enabling our engineering teams to ship fast. You'll... ...automation and tooling that improves reliability and partnering with... ...scale with the product Manage and optimize our compute, networking... ...'ll Bring ~5+ years in site reliability engineering,...Full timeWork at officeWork from home$159.2k - $301.6k
...APIs for two primary uses: (1) managing Graph plugins, their... ...Graphs on the cloud. In this reliability-focused role, you will own the... ...'ll partner with the backend engineers building these APIs to make sure... ...~5-10 years of experience in site reliability engineering, infrastructure...Temporary workLocal areaWorldwide$287k
...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to... ...hit our SLAs. What? We're looking for a Senior or Staff Site level Reliability Engineer as part of Infrastructure team to: Own uptime...Contract workWork at officeRemote work$220k - $235k
...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts... ...Wave and Gartner Magic Quadrant for Contract Lifecycle Management, a Fortune Great Place to Work, and one of Fast Company's...Full timeContract workWork at office$200k - $260k
...Infrastructure Team as a technical leader driving reliability, automation, and scalability across the... ...practices across teams, mentor senior engineers, and be a primary escalation point for... ...engineers, without needing formal management authority to do it ~ Strong bias for...Casual workWork at officeRemote workFlexible hours$194k - $267k
...automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses...Permanent employmentWork at officeLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager of Site Reliability Engineering. Be the first to apply!
- principal cloud engineer San Francisco, CA
- senior principal engineer San Francisco, CA
- assistant chief engineer San Francisco, CA
- principal infrastructure engineer San Francisco, CA
- general engineer San Francisco, CA
- principal safety engineer San Francisco, CA
- principal application developer San Francisco, CA
- principal engineer San Francisco, CA
- director of product engineering San Francisco, CA
- director data engineering San Francisco, CA

