Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineering

$204k - $306k
Full-time

Okta

Secure Every Identity, from AI to Human

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

Manager, Site Reliability Engineering

San Francisco, California

Secure Every Identity, from AI to Human

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

**This position requires 2 days a week in our San Francisco Office.

The IDaaS Site Reliability Engineering Group

Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling.

As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling.

What you’ll be doing

  • Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform.
  • Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization.
  • Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns.
  • Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.
  • Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.
  • Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management
  • Manage service and business expectations and prioritize resource allocation
  • Maintain a deep knowledge of industry best practices, evolving trends, and technologies

What you’ll bring to the role

  • 3+ years of experience in technical leadership & people management
  • Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scale
  • Experience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus.
  • Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines
  • Strong background and hands-on experience in SW development, PaaS and automation
  • Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.
  • Effective verbal, written communication and interpersonal skills
  • Computer Science Degree or related degree or equivalent experience

Additional requirements:

  • This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.

#LI-Hybrid

P24518_3462184

Below is the annual base salary range for candidates located in San Francisco Bay Area. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit:

The annual base salary range for this position for candidates located in the San Francisco Bay area is between:

$204,000—$306,000 USD

The Okta Experience

  • Supporting Your Well-Being
  • Driving Social Impact
  • Developing Talent and Fostering Connection + Community

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.

If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation.

Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.
Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineering in Washington DC vacancy
  • $166k - $220k

     ...expectations. Our systems integration engineers internalize the nuances of each deployment...  ...ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly...  ...the developer experience. You will be managing cloud deployments in AWS, Azure and on... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    1 day ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient...  ...:Infrastructure & Automation• Design, deploy, and manage cloud infrastructure using Infrastructure as Code (IaC) principles... 
    Suggested
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    4 days ago
  • $115.5k - $164.8k

     ...company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will...  ...previously required human intervention with reliable, tested automation. You will also...  ...years of applicable experience.Experience managing cloud platforms such as Azure, AWS, or similar... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    4 days ago
  • $150k - $180k

     ...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale...  ...organization.This position is based on-site in either our Arlington, VA office,...  ...closely with cross-functional teams, product managers, and stakeholders to align on technical... 
    Suggested
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    2 days ago
  • $165k - $270k

     ...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology...  .... RESPONSIBILITIES: Develop automation to deploy and manage compute resources both on-premises and in the cloudDeploy and... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    4 days ago
  • $125k - $185k

     ...children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-...  ...and observability & systems toolsComfort with configuration management, load balancing, monitoring & alerting infrastructure, and... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    2 days ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of...  ...maintain highly available production systems.Define and manage SLIs, SLOs, and error budgets.Automate operational tasks and... 
    Remote work

    Govcio

    Arlington, VA
    2 days ago
  • $166k - $220k

     ...expectations. Our systems integration engineers internalize the nuances of each deployment...  ...ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly...  ...the developer experience. You will be managing cloud deployments in AWS, Azure and on... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    4 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the...  ..., and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate...  ...& systems toolsComfort with managing large scale production systems and technologies... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    4 days ago
  • $112k - $179k

     ...The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington...  ...automation solutions for infrastructure and application management.Drive modernization efforts by engineering and implementing... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    2 days ago
  • $174k - $238k

     ...The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products...  ...balancing, ingress, TLS, service networking, and traffic management.Experience with observability platforms, monitoring strategies... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    13 hours ago
  • $207k - $284.9k

     ...all in on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    4 days ago
  • $174k - $239k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for...  ...networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services,... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    1 day ago
  • $95k - $171k

     ...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's...  ...inference platform. As an Site Reliability Engineer II, you will be responsible for:...  ...integrating into Akamai's existing incident management processes Contributing to SLO tracking... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    5 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA....  ...Remote unavailable. Modality: On‑Site only. Must live within...  ...scalability, performance, and reliability across environments. What You...  ...powers AI workloads at scale Manage and automate GPU compute clusters... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    14 hours ago
  • $182k - $250.8k

     ...at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    3 days ago
  •  ...the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-...  ...technical ownership and customer-facing engineering: you'll define how we measure...  ...within the restricted environment. Manage containerized services (Docker,... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Inc.

    Arlington, VA
    1 day ago
  •  ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development and Systems engineering...  ...life cycle and mitigates vulnerabilities. Monitors and manages the Stability, Availability, and Performance of enterprise... 
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    2 days ago
  • $106.3k - $221.1k

     ...missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability...  ...-D # GCSA # GISF # SSCP ~8 years of experience managing reliability, uptime, and automating operations for large... 
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    2 days ago
  • $168k - $200k

     ...What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront...  ...experience , including workspace setup, cluster/job management, and integration with CI/CD and data orchestration tools.... 
    Remote work

    Datavant

    Washington DC
    1 day ago
  • $121.4k - $218.6k

     ...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure...  ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for:...  ...rotations, spearheading real-time incident management, and managing high-severity service... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    5 days ago
  •  ...plan Paid maternity leave 401(k) Get notified when a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle, WA $115,000.00-$175,000.00 5 months ago Senior ServiceNow... 
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    14 hours ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology...  ...term reliability of cloud native systems. Our SREs don’t just manage infrastructure; they build it using Infrastructure as Code... 
    Remote work

    Noctua Technology

    Washington DC
    3 days ago
  • $160k - $180k

     ...Site Reliability Engineer Location: Hybrid – Washington DC/Virginia/Maryland metro with the ability to travel to Patuxent River, MD, as needed...  ...as CI/CD pipelines, log/application monitoring, cluster management, and configuration management Migrate application... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Fortress Information Security

    Washington DC
    5 days ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The...  ...Responsibilities Key Responsibilities Capacity Ingestion and Management: -Takes proactive steps to design and architect... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    3 days ago
  • $165k - $270k

     ...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX's Starlink...  .... RESPONSIBILITIES: Develop automation to deploy and manage compute resources both on-premises and in the cloud Deploy... 
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    4 days ago
  • $94k - $142.3k

     ...customer and partner enablement, applications engineering, infrastructure, collaboration,...  ...delivered globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of...  ...’ll Actually Be Doing...Respond to and manage major incidents affecting internal... 
    Full time
    Shift work

    Salesforce

    Washington DC
    2 days ago
  • Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex...

    Okta, Inc.

    Washington DC
    3 days ago
  • $105.79k - $141.05k

     ...networking at scale. As Lead SRE, you'll own the reliability of that platform — partnering with...  ...accountable for platform observability, incident management, and automation, and you'll coordinate across architecture, engineering, and systems development organizations to... 
    Temporary work
    Remote work

    Lumen Inc

    Washington DC
    1 day ago
  • $169.3k - $304.7k

     ...team! Our team designs, develops, and manages applications and infrastructure that...  ...fast, efficient, scalable, and reliable routing software and infrastructure that...  ...our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineering. Be the first to apply!