Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineering

$204k - $306k

Okta

Secure Every Identity, from AI to Human

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Manager, Site Reliability Engineering

San Francisco, California

Secure Every Identity, from AI to Human

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

**This position requires 2 days a week in our San Francisco Office. 

The IDaaS Site Reliability Engineering Group

Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling. 

As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling. 

What you’ll be doing 

  • Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform.
  • Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization.
  • Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns.
  • Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.
  • Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.
  • Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management 
  • Manage service and business expectations and prioritize resource allocation
  • Maintain a deep knowledge of industry best practices, evolving trends, and technologies

What you’ll bring to the role

  • 3+ years of experience in technical leadership & people management 
  • Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale
  • Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus.
  • Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines
  • Strong background and hands-on experience in SW development, PaaS and automation
  • Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.
  • Effective verbal, written communication and interpersonal skills
  • Computer Science Degree or related degree or equivalent experience 

Additional requirements:

  • This position requires the ability to access federal environments and/or have access to protected federal data.  As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.

#LI-Hybrid

P24518_3462184

Below is the annual base salary range for candidates located in San Francisco Bay Area. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: .   

The annual base salary range for this position for candidates located in the San Francisco Bay area is between: $204,000—$306,000 USD

The Okta Experience

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.

If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please  use this Form to request an accommodation.

Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please  click here to view our full NYC AEDT Notice.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineering in Washington DC vacancy
  • $151k - $297k

     ..., you will partner with SRE leaders and engineers to scale the platform that underpins all...  ...program execution, strengthen production reliability practices, and coordinate cross-...  ...criteria with SRE engineers and leaders. Manage dependencies across platform teams, keep... 
    Suggested
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Washington DC
    4 days ago
  • $165k - $230k

     ...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology...  .... RESPONSIBILITIES: Develop automation to deploy and manage compute resources both on-premises and in the cloudDeploy and... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    7 days ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of...  ...maintain highly available production systems.Define and manage SLIs, SLOs, and error budgets.Automate operational tasks and... 
    Suggested
    Remote work

    Govcio

    Arlington, VA
    5 days ago
  • $125k - $185k

     ...children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-...  ...and observability & systems toolsComfort with configuration management, load balancing, monitoring & alerting infrastructure, and... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    5 days ago
  • $185k - $230k

    As a Sr. Site Reliability Engineer (SRE) III, you’ll work as part of a collaborative and high-performing team providing your expertise to deliver...  ...automated CI/CD pipelines, monitoring, and configuration management workflows to support reliable software delivery and... 
    Suggested
    Full time
    Local area
    Immediate start

    MetroStar Systems

    Washington DC
    5 days ago
  • $166k - $220k

     ...expectations. Our systems integration engineers internalize the nuances of each deployment...  ...ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly...  ...the developer experience. You will be managing cloud deployments in AWS, Azure and on... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    4 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient...  ...:Infrastructure & Automation• Design, deploy, and manage cloud infrastructure using Infrastructure as Code (IaC) principles... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    7 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the...  ..., and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate...  ...& systems toolsComfort with managing large scale production systems and technologies... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    7 days ago
  • $112k - $179k

     ...The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington...  ...automation solutions for infrastructure and application management.Drive modernization efforts by engineering and implementing... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    5 days ago
  • OB SUMMARYThe Systems Engineer - Site Reliability Engineering (SRE) is responsible for the reliability, scalability, and performance of mission...  ...globally. This role involves overseeing incident management, driving automation efforts, and working closely with cross... 
    Full time
    For contractors
    Work at office
    Remote work
    Flexible hours
    Shift work

    Marriott International

    Bethesda, MD
    4 days ago
  • $150k - $180k

     ...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale...  ...organization.This position is based on-site in either our Arlington, VA office,...  ...closely with cross-functional teams, product managers, and stakeholders to align on technical... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    5 days ago
  • $115.5k - $164.8k

     ...company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will...  ...previously required human intervention with reliable, tested automation. You will also...  ...years of applicable experience.Experience managing cloud platforms such as Azure, AWS, or similar... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    7 days ago
  •  ...some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for...  ...or equivalent practical experience.Proven experience managing SaaS or PaaS systems at enterprise scale (multi-region, multi... 
    Temporary work

    Kong

    Washington DC
    5 days ago
  • $174k - $239k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for...  ...networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services,... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    4 days ago
  • $174k - $238k

     ...The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products...  ...balancing, ingress, TLS, service networking, and traffic management.Experience with observability platforms, monitoring strategies... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  • $182k - $250.8k

     ...at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    5 days ago
  • $207k - $284.9k

     ...all in on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    7 days ago
  •  ...Site Reliability Engineer ValidaTek is building teams of Site Reliability Engineers (SRE's) to support internal and external engineering and...  ...and automation System provisioning and lifecycle management; experience with Red Hat Satellite Networking Container... 

    ClearanceJobs

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable...  ...applications. ~ Experience in Docker orchestration and management. ~ Experience with Kubernetes. Education: ~... 
    Work experience placement

    Samprasoft

    Washington DC
    5 days ago
  •  ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington...  ...-as-Code (IaC): Automate the provisioning and management of cloud infrastructure using IaC tools like Terraform, CloudFormation... 

    Software Technology Inc

    Washington DC
    3 days ago
  • $75.7k - $136.3k

     ...Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track Key Performance Indicators (KPIs) and Service Level Objectives (SLOs), identify and resolve performance bottlenecks, and perform... 
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    2 days ago
  •  ...Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian...  ...1. Reliability & Performance Engineering SLA/SLO Management: Define, monitor, and maintain Service Level Objectives (... 
    Local area

    Tiger Analytics

    Washington DC
    2 days ago
  •  ...and In-Q-Tel. Mission | On Site | Full Time | Active TS/SCI...  ...customer site, ensuring the reliability and performance of Twenty's...  ...ownership and customer-facing engineering: you'll define how we measure...  ...the restricted environment. Manage containerized services (... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Technologies

    Arlington, VA
    2 days ago
  •  ...TENEX Staff Site Reliability Engineer TENEX is an AI-native, automation-first, built-for-scale Managed Detection and Response (MDR) provider. We are a force multiplier for defenders, helping organizations enhance their cybersecurity posture through advanced threat detection... 
    Work from home

    TenEx

    Washington DC
    3 days ago
  •  ...Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical...  ...automated CI/CD pipelines, monitoring, and configuration management workflows across all environments. Provision, configure... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    18 days ago
  • $112k - $218.4k

     ...: Our team is looking for a Senior Active Directory Site Reliability Engineer. Our mission is to improve the availability, latency, performance...  ...Access Protocol (LDAP), Kerberos, Windows New Technology LAN Manager (NTLM), Group Policy, Windows Time, Trusts, Delegation,... 
    Full time
    Local area

    Microsoft

    Washington DC
    6 days ago
  •  ...Job Description Job Description Site Reliability Engineer II Metro DC · Hybrid · 24/7 FedRAMP Operations · Rotational Shift · Initial Contract till March 27. KEY REQUIREMENT This role requires US citizenship and residence on US soil. It sits within a FedRAMP... 
    Hourly pay
    Contract work
    For contractors
    Shift work
    Night shift
    Weekend work

    C-Serv

    Arlington, VA
    10 days ago
  • $160k - $210k

     ...Dynamic Deals through their preferred DSP, leveraging our managed service DSP, or utilizing our industry-first ContextGPT product...  .... Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management... 
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Washington DC
    10 days ago
  • $175k - $250k

    Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote...  ...unavailable. Modality: On‑Site only. Must live within...  ...scalability, performance, and reliability across environments. What You...  ...powers AI workloads at scale Manage and automate GPU compute clusters... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineering. Be the first to apply!