Senior Manager, Site Reliability Engineering - Infrastructure Platform
$266k - $398kOkta, Inc.
Director, Site Reliability Engineering – Infrastructure Platform Okta is The World’s Identity Company. Okta provides secure access, authentication, and automation, placing identity at the core of business security and growth. The Infrastructure Platform and Shared Services Team Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on Amazon Web Services (AWS) across multiple availability zones and geographically separated regions. The service is designed for high throughput, and 99.999% availability. We’re looking for a technical leader to help us to continue to scale the service with great people and reliable, cost-effective and efficient infrastructure, processes and tooling. What You’ll Be Doing Lead the infra platform and shared services org and various initiatives across SRE & Infrastructure organization. Lead the DevOps transformation, microservice journey, and next generation infra platform capabilities in partnership with architects and product engineering. Build a world‑class observability platform and monitoring capabilities enabled with self‑service. Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and intuitive self‑service capabilities. Own the design and operation of scalable, self‑service Cloud infrastructure platforms (e.g., Kubernetes, service mesh, CI/CD pipelines, IaC & Edge Infrastructure). Lead, mentor, and grow a high‑performing team of engineers and managers across platform, infrastructure, and shared services domains. Perform engineering design evaluations and ensure the completion of projects within resource, budget, and scheduling constraints. Improve SDLC processes for Cloud infrastructure as code, including the maturity of CI/CD pipelines, change and release management. Manage service and business expectations and prioritize resource allocation. Maintain a deep knowledge of industry best practices, evolving trends, and technologies. What You’ll Bring To The Role 8+ years of experience in technical leadership & people management. Extensive experience using Agile and DevOps methodologies to build product infrastructure and shared services at scale. 3+ years of experience running large‑scale infrastructure platforms supporting a SaaS/Cloud service in a public Cloud, preferably AWS. Experience supporting a multi‑Cloud environment will be a plus. Strong expertise in cloud‑native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines. Strong background and hands‑on experience in SW development, PaaS and automation. Deep experience with building and operating observability platforms and monitoring tools (Grafana, Splunk, APM, etc.) in a large‑scale environment. Demonstrated ability to lead cross‑functional teams and manage large‑scale programs. Effective verbal, written communication and interpersonal skills. Computer Science Degree or related degree or equivalent experience. Additional Requirements This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g., a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee.22CFR120.15) upon hire. Compensation and Benefits Annual base salary range for candidates located in California: $266,000—$398,000 USD. Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave). To learn more about our Total Rewards program please visit: End‑of‑Job Legal Statements Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. #J-18808-Ljbffr
- ...the trusted, neutral infrastructure that enables organizations... .... The Infrastructure Platform and Shared Services... ...great people and reliable, cost-effective, and... ...tooling. As the Sr. Manager of Infrastructure Platform... ...and product engineering Build a world-class observability...Senior
- ...affordable. Join us in building a platform that empowers innovators... ...Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional... ...incident response systems, managing capacity across our distributed...Senior
- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public... ...systems, design and improve the infrastructure behind our production environments,... ...hand experience with configuration management and infrastructure as code (Ansible...Senior
- ...meaningful impact. As a Senior Lead Infrastructure Engineer at JPMorgan Chase within... ...requirements Works with other platforms to architect and implement... ...and the ability to manage assigned tasks efficiently... ...health care coverage, on-site health and wellness centers...Senior
- ...Serval is the AI platform for IT teams — replacing... ...automation engine adopted by HR, Finance... ...offboarding, software access management, and the long tail of... ...a Software Engineer, Infrastructure, you'll build and... ...availability, performance, and reliability of production systems...SeniorFull time
- ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our... ...industry leading hybrid cloud infrastructure, innovative data capabilities with leverage... ...networking, kernel drivers, package management, etc. or have a good idea where to start...SeniorTemporary workWork experience placement
$230k - $265k
...tools they need through the platforms they already sell on.... ...flexible funding, spend management, and savings tools to their... ...an experienced software engineer to join our Infrastructure team and help build the foundation... ...and services, ensuring reliability, scalability, and...SeniorFull timeWork from homeFlexible hours$170k - $220k
..., and operated. Our platform gives engineers real-time observability... ...faster, more reliable development. Sift was... ...reliability demanded new infrastructure. Founded by a team... ...infrastructure management using modern DevOps... ...company, being nearby for site visits and...SeniorFull timeRelocation$130k - $175k
...groundbreaking educational platform that promotes student... ...curriculum management functionality, Kiddom... ...and empower Kiddom’s engineering by building a scalable... ...services. Practicing Infrastructure as Code (IaC) wherever... ..., past experience, seniority, and demonstrated role...SeniorPermanent employmentFull timeWork at officeLocal areaRemote work$210.8k - $272.8k
...their homes. About the Site Reliability Engineering Team The Site Reliability... ...reliable, secure, and scalable platform vital for a seamless user... ..., and resiliency of infrastructure and software, working with... ..., and infrastructure management Troubleshoot and debug critical...SeniorLocal area$175k - $250k
...0.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco... .... Modality: On-Site only. Must live within... ...flexibility. Their platform allows users to connect... ..., performance, and reliability across environments.... ...workloads at scale Manage and automate GPU...SeniorFull timeRemote workRelocationRelocation package$140k - $225k
...world-class team—engineers, designers, researchers... ...Role As the Senior Software Engineer... ...and Development Infrastructure, you will play a... ...the quality and reliability of our AI... ...on experience in platforms like Jenkins, GitHub... ...proven ability to manage cross-functional...SeniorFull timeTemporary workLocal areaFlexible hours$210k - $240k
Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies... ...SRE) to help us scale our platform with reliability,... ...automate, and maintain the infrastructure that powers our core platform... ...rollback mechanisms, and config management Implement and maintain...SeniorFull time$190k - $270k
...AI Chopping Block, Inc. is seeking an AI Infrastructure Engineer in San Francisco, CA, to ensure the optimal operation of user-facing services and production systems. This role requires expertise in building infrastructure with Ansible, Terraform, and Kubernetes, along...Senior$186.07k - $218.9k
...Senior Software Engineer, Developer Infrastructure What you’ll be doing (job duties) Design, build, and operate core... ...infrastructure — that are fast, reliable, and secure by default at Coinbase... ...understand pain points and turn them into platform improvements. Lead technical...SeniorLocal areaRemote work- ...A tech company specializing in AI infrastructure is seeking a Software Engineer to build a scalable compute platform for its generative video models. The ideal candidate... ...of experience in MLOps or AI infrastructure management, along with strong Python and C++ skills. This...Senior
- ...Senior Software Engineer, ML Infrastructure Posted 2 days ago. Be among the first 25 applicants. Get AI‑powered... ...About LMArena: LMArena is the open platform for evaluating how AI models perform... ...to understand real‑world reliability, alignment, and impact. Our leaderboards...SeniorPermanent employmentFull timeWork at office
$320k
...mission is to create reliable, interpretable, and steerable... ...researchers, engineers, policy experts, and... ...We're building the infrastructure that enables Claude to... ...code, calling APIs, managing files, and completing... ..., infrastructure, or platform services at the hyper...SeniorFull timeWork at officeImmediate startVisa sponsorshipFlexible hours$190k - $221k
...Senior Software Engineer, Infrastructure TRM Labs is a blockchain intelligence company committed to fighting... ...scientists, engineers, and product managers to design and build scalable infrastructure... ...to understand the availability, reliability, and sustainability of our...SeniorRemote work- A leading AI research firm in San Francisco seeks a Staff Infrastructure Engineer to identify and resolve infrastructure bottlenecks and design large-scale systems for AI training. The ideal candidate has over 3 years of experience in infrastructure engineering and strong...Senior
$155k - $216k
...learning. As the leading identity platform for education, more than 111,000... ...visit us at clever.com anytime. The Infrastructure team handles platform engineering at Clever. It builds and evolves... ...and safely, while meeting reliability, security, and compliance expectations...SeniorWork at officeWorldwideFlexible hoursShift work$200k - $250k
...Role Overview As a Software Engineer on the Core Infrastructure team at Harvey, you will... ...with Harvey’s AI platform, processing billions of prompt... ...operational excellence, ensuring reliability, scalability, and... ...networking, and container management Lead technical initiatives...SeniorInternshipWorldwideRelocation package- ...meaningful impact. As a Senior Lead Software Engineer - Windows Server... ...Corporate Sector Compute Infrastructure Platform (CIP) organization, you exhibit... ...technology processes Manages software engineering in accordance... ...health care coverage, on-site health and wellness...Senior
- ...delivering an AI-powered platform that governs and... ...As a Staff Platform Engineer, you will play a critical... ...role. You will own reliability for major platform domains... ...the shared infrastructure services and platforms... ...Architect, implement, and manage highly available and...Senior
- ...Founding DevOps Engineer role at Retell... ..., and ship reliable releases that... ...into reusable platform capabilities.... ...tenant concerns. Manage cloud deployments... ...with on-prem infrastructure teams;... ...Deploy at customer sites (cloud or on-prem... .... Seniority level Seniority...SeniorFull timeWork at officeLocal areaWork from homeVisa sponsorshipFlexible hours
$150k - $200k
...workload segmentation as our infrastructure footprint grows Owning our... ...evolving our vulnerability management, misconfiguration detection... ...— Terraform reviews, platform backlog, on‑call rotation... ...6+ years of hands‑on cloud engineering experience, with substantial...SeniorWork at officeRemote work- Grow Therapy in San Francisco is seeking a Senior AI Enablement Engineer to define how AI transforms operations across the organization. You will design and build foundational AI infrastructure that enhances efficiency. Responsibilities include implementing AI systems,...SeniorFlexible hours3 days per week
$180k - $220k
...Francisco, CA · On-site · Full-time Compensation... ...out its founding engineering team. Founded 2025... ...AI research / data infrastructure The Role As a Senior Software Engineer, Infrastructure and Platform, you'll design and... ...with strong reliability guarantees Create...SeniorFull timeVisa sponsorship$180k - $270k
...guidance to teams across engineering, product, and... ...and machine learning infrastructure to enable Plaid engineers... ...developing on top of this platform extremely simple for... ..., and project management skills, with proven ability... ...design, ensuring reliable and scalable solutions...SeniorWork experience placementLocal area- ...designs, builds, and operates critical infrastructure that enables research at OpenAI. Our... ...size of our workloads, while remaining reliable and easy to use. About the Role We... ...’re looking for a staff-level software engineer to own production-critical...Full timeWork at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Manager, Site Reliability Engineering - Infrastructure Platform. Be the first to apply!
- site reliability engineer sre San Francisco, CA
- site reliability engineer San Francisco, CA
- IT operations director San Francisco, CA
- senior infrastructure engineer San Francisco, CA
- infrastructure engineering manager San Francisco, CA
- infrastructure engineer San Francisco, CA
- remote infrastructure engineer San Francisco, CA
- security infrastructure engineer San Francisco, CA
- infrastructure developer San Francisco, CA
- lead infrastructure engineer San Francisco, CA



