Manager, Site Reliability Engineering
$204k - $306kOkta
Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from AI to HumanIdentity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.**This position requires 2 days a week in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput and 99.999 availability. We're looking for a technical leader to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and tooling. As the Manager of Infrastructure Platform and Shared Services, you will oversee multiple teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling. What you’ll be doing Managing a team of SRE’s supporting various workloads and teams that support our IDaaS platform.Drive the microservice journey, DevOps maturity, and workload reliability in tandem with architects and teams across the organization.Accelerate the velocity of SRE and product engineering by developing powerful tooling, intuitive self-service capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing team of engineers and managers across platform, infrastructure, and shared services domains.Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints.Improve SDLC processes for Cloud infrastructure as a code, including the maturity of CI/CD pipelines, change and release management Manage service and business expectations and prioritize resource allocationMaintain a deep knowledge of industry best practices, evolving trends, and technologiesWhat you’ll bring to the role3+ years of experience in technical leadership & people management Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared service at scaleExperience running large-scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi-Cloud environment will be a plus.Strong expertise in cloud-native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelinesStrong background and hands-on experience in SW development, PaaS and automationDeep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM etc.) in a large scale environment.Effective verbal, written communication and interpersonal skillsComputer Science Degree or related degree or equivalent experience Additional requirements:This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.#LI-HybridP24518_3462184Below is the annual base salary range for candidates located in San Francisco Bay Area. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: The annual base salary range for this position for candidates located in the San Francisco Bay area is between: $204,000—$306,000 USDThe Okta ExperienceSupporting Your Well-BeingDriving Social Impact Developing Talent and Fostering Connection + CommunityWe are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation.Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice.
$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ....Required Skills & Experience (The Essentials)Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at...SuggestedPermanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $267k
..., automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses...SuggestedPermanent employmentWork at officeLocal areaWorldwideFlexible hours- ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend... ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable...Suggested
$232k - $319k
...scale the service with great people and reliable, cost-effective, and efficient infrastructure... ..., processes, and tooling. As the Sr. Manager of Infrastructure Platform and Shared... ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful...SuggestedPermanent employmentLocal areaWorldwideFlexible hours$102.1k - $202.2k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole... ...EngineeringDiscipline: Site Reliability EngineeringCompany:... ...secure, resilient, and easier to manage for customers operating in highly... ..., you will collaborate with engineers across disciplines to deliver...SuggestedOngoing contractWork experience placementLocal areaRemote work3 days per week$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical... ...and alerting systems.The Fleet Management team provides the core runtime environment... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...Work at officeLocal areaRemote workWorldwideFlexible hours$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on... ...cloud-native distributed systems.Employ technical project management skills, with the ability to properly scope, plan, and define...Work experience placementWork at officeRemote workFlexible hours$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole... ...EngineeringDiscipline: Site Reliability EngineeringCompany:... ...As a Senior Site Reliability Engineer, you will lead reliability improvements... ...eligibility requirements.For manager-level roles, a Tier 5 (T5)...Ongoing contractLocal area3 days per week- ...and private sectors, and the Department of Defense. Our services span all aspects of business, providing a holistic approach for managing an organization.Job DescriptionResponsible for the monitoring, provisioning, resiliency, and customer interactions. Working extensively...
$151.2k - $204.6k
Would you like to be an engineer who builds the systems that power advertising at scale,... ...advertising queries every day, where latency, reliability, and quality translate directly into... ...Development Engineer, operating as a Site Reliability Engineer, to raise the reliability...Flexible hours$143k - $194k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...Full timeTemporary workWork experience placementImmediate start$165k - $270k
...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology... .... RESPONSIBILITIES: Develop automation to deploy and manage compute resources both on-premises and in the cloudDeploy and...Permanent employmentTemporary workWork at officeImmediate startMonday to FridayWeekend work$134.25k - $214.8k
...where you matter.Your ImpactAre you an engineer who gets excited about the challenge... ...the Observability team within Axon's Site Reliability organization — a focused team responsible... ...ArgoCD, and Helm — including capacity management, cybersecurity requirements and...Work experience placementWork at officeRemote work- ...enterprise solutions within the Information Security team. The successful candidate will apply an engineering-based approach to solving complex security and reliability challenges, leveraging machine data analytics and log analysis across the enterprise. Designing systems...Full timeInternshipSummer internshipWork at officeLocal areaRemote workFlexible hours
$165k - $270k
...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building... ...infrastructure to support a multi-region environment Manage petabyte scale bare metal compute clusters Closely collaborate...Permanent employmentTemporary workWorldwideWeekend work$165k - $230k
...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging... ...alerting infrastructure to support a multi-region environment Manage petabyte scale bare metal compute clusters Closely...Permanent employmentTemporary workWorldwideWeekend work- ...day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance... ...individual-contributor role, not a management track. You'll spend your time in the... ...years of hands-on Cloud Operations and Site Reliability Engineering, operating...Full time
- ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry... ...clear, actionable signals into system health. Incident Management: Act as an incident commander during major outages, lead...Full timeWork at office
- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems... ...and observability solutions Manage incident response and post-mortems Improve... ...to have Experience with chaos engineering Knowledge of distributed systems...Remote workFlexible hours
- ...limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise... ...transaction processing and asset management. We offer a competitive total rewards package... ...comprehensive health care coverage, on-site health and wellness centers, a...
$55k - $151.47k
...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in... ...data integrity and accessibility- Leading incident management and resolution efforts to maintain operational continuityWhat...Full timeH1b- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base... ...collaborative workspaces, Mail’s inbox management, and Go, the proactive AI assistant that... ...for building software to ensure the reliability of our back-end systems, working with engineers...WorldwideHome officeFlexible hours
$95k - $134k
...Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work... ...closely with peer teams, platform/software architects and management to drive key reliability improvements. Willingness to...Temporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours- .... We are responsible for the reliability of all the company's major data... ..., services, and query engines. We serve business needs across... ...utilized effectively.- Incident Management: Lead efforts to troubleshoot... ...emerging technologies related to site reliability and...
- ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health...
- ...offer IT solutions across the disciplines of program/project management, applications development, infrastructure, Cyber security,... ...(ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months...Local areaWorldwide
$94k - $142.3k
...customer and partner enablement, applications engineering, infrastructure, collaboration,... ...delivered globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of... ...’ll Actually Be Doing...Respond to and manage major incidents affecting internal...Full timeShift work- ...what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology... ...transaction processing and asset management. We offer a competitive total rewards... ...comprehensive health care coverage, on-site health and wellness centers, a retirement...
$145k - $160k
...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives... ...Active-Active) architectures Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics...Temporary workRemote workFlexible hours- ...what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology... ...transaction processing and asset management. We offer a competitive total rewards... ...comprehensive health care coverage, on-site health and wellness centers, a retirement...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager, Site Reliability Engineering. Be the first to apply!
- official site Bellevue, WA
- site services specialist Bellevue, WA
- construction site safety Bellevue, WA
- IT site lead Bellevue, WA
- site leader Bellevue, WA
- site safety Bellevue, WA
- historic site Bellevue, WA
- junior website developer Bellevue, WA
- on-site clinical research associate (traveling/remote) Bellevue, WA
- site reliability engineering manager



