Staff Site Reliability Engineer - Kubernetes

$194k - $267k

Okta, Inc.

Secure Every Identity, from AI to Human

Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.

This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.

Workforce Identity Cloud

Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities-like reducing costs, and doing more for your customers.

If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, "If you have to do something more than once, automate it" and who can rapidly self-educate on new concepts and tools.

Position Overview:

The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh.

Key Responsibilities:

Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency.
AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement best practices for cost management, scaling, and security within AWS.
Helm Management: Utilize Helm to automate and streamline the deployment of applications and services to Kubernetes clusters. Create, maintain, and manage Helm charts for production-ready deployments.
Karpenter Implementation: Implement and manage Karpenter to dynamically scale Kubernetes clusters in response to workload demands.
Istio Service Mesh Management: Configure and manage Istio to provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained traffic management, service discovery, and policy enforcement.
Platform Automation & Scaling: Automate the deployment, scaling, and management of infrastructure and applications. Work with CI/CD pipelines to ensure a seamless flow from development to production with minimal downtime.
Incident Management & Troubleshooting: Respond to incidents, troubleshoot, and resolve system issues related to performance, availability, and security in a timely and effective manner.
Security & Compliance: Design and implement secure cloud infrastructure with appropriate access controls, network security, and compliance frameworks.
Documentation & Knowledge Sharing: Create and maintain detailed documentation for Kubernetes platform setup, operational procedures, and best practices. Promote knowledge sharing across teams.

Required Qualifications:

4+ years of experience with Kubernetes/Helm;
4+ years of Experience with Terraform.
5+ years of Experience with AWS
Experience with multi-region cloud environments.
Proven experience with AWS (EC2, RDS, S3, CloudFormation, IAM, etc.) and solid understanding of cloud-native architectures.
Strong expertise in Kubernetes platform creation, management, and optimisation (e.g., setting up highly available clusters, networking, and storage).
Hands-on experience with Helm for Kubernetes application deployment and management.
Practical experience with Karpenter for dynamic scaling of Kubernetes clusters and optimising resource usage.
Expertise in managing and securing Istio for service mesh, including traffic management, security, and observability features.
Proficiency in CI/CD pipelines and automation tools (e.g., Jenkins, GitLab, CircleCI, Terraform, Ansible, Spinnaker).
Strong scripting and automation skills in Python, Bash, or Go for infrastructure management and platform automation.
Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack.

Preferred Qualifications:

Understanding of security best practices for cloud platforms and Kubernetes (e.g., role-based access control (RBAC), encryption, and compliance frameworks).
Familiarity with Docker and containerization principles.
Bachelor's degree in Computer Science, Engineering, or related field (or equivalent professional experience).
Certifications (Preferred): CKA (Certified Kubernetes Administrator), CKAD (Certified Kubernetes Application Developer), or AWS Certified DevOps Engineer are highly desirable.

Additional requirements:

This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.
Requires in-person onboarding and travel to our San Francisco, CA HQ office or our Chicago office during the first week of employment.

#LI-Hybrid

#LI-LSS1

requisition ID- (P16373_3396241)

The annual base salary range for this position for candidates located in the San Francisco Bay area is between: $194,000—$267,000 USD

Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit:

The annual base salary range for this position for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York, and Washington is between: $174,000—$214,000 USD

The Okta Experience

Supporting Your Well-Being
Driving Social Impact
Developing Talent and Fostering Connection + Community

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws.

If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding pleaseuse this Form to request an accommodation.

Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, pleaseclick here to view our full NYC AEDT Notice.

Apply

Vacancy posted 4 days ago

Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer - Kubernetes in San Francisco, CA vacancy

Senior Site Reliability Engineer: Cloud & Kubernetes Lead
...leading technology firm in San Francisco is seeking a Senior Site Reliability Engineer to maintain and improve cloud infrastructure. The ideal... ...experience as an SRE or DevOps engineer and strong expertise in Kubernetes. This role focuses on automation and enhancing system...
Suggested
TechChain Talent
San Francisco, CA
3 days ago
Site Reliability Engineer
...our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as... ...dive into unfamiliar backend codebases ~ Strong Go and Kubernetes experience. ~ Familiarity with observability and monitoring...
Suggested
Worldwide
Shift work
Happy Robot
San Francisco, CA
4 days ago
Site Reliability Engineer
...Open Source LLM Gateway Engineer LiteLLM is an open-source LLM Gateway with 34K+ stars... ...our 6th Engineer focused on owning reliability, performance, and infrastructure stability... ...FastAPI, Redis, Postgres, Prisma ORM, Kubernetes, Prometheus, Docker. Who we are looking...
Suggested
BerriAI
San Francisco, CA
3 days ago
Site Reliability Engineer
...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...: Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems,...
Suggested
Flexible hours
Baseten
San Francisco, CA
2 days ago
Site Reliability Engineer
...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure... ..., SR-IOV, high-throughput networking) ~ Experience with Kubernetes or similar orchestrators ~ Familiarity with...
Suggested
Blaxel, Inc
San Francisco, CA
5 days ago
Site Reliability Engineer
...Francisco, NYC, or London offices. About the Role As a Site Reliability Engineer (SRE) at Mercor, you'll own production reliability across... ...scratch. Hands-on experience in the AWS ecosystem, Kubernetes, and modern IaC tooling (Terraform, Spacelift, etc.). Benefits...
Work at office
Relocation package
Mercor Alabaster
San Francisco, CA
2 days ago
Senior Site Reliability Engineer, Fleet Management
$127k - $249k
THE TEAM Platform Engineering is the department within SRE that is responsible for a range... ...these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...
Work at office
Local area
Remote work
Worldwide
Flexible hours
MongoDB
San Francisco, CA
4 days ago
Site Reliability Engineer (SRE)
$350k
...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity through advancing collaborative general... ...for long-running distributed jobs. Expertise in Kubernetes at scale: deploying, operating, debugging, and tuning clusters...
Local area
Visa sponsorship
Work visa
Relocation package
Thinking Machines Lab
San Francisco, CA
3 days ago
Senior Site Reliability Engineer
...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with... ...senior engineering candidate to join the Site Reliability organization in San Francisco. Working... ...containerized architectures (Docker, Kubernetes) and orchestration platforms. ~...
Worldwide
Weekend work
Salesforce
San Francisco, CA
2 days ago
Site Reliability Engineer
...DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you... ...Systems Design: Experience with container orchestration (Kubernetes) and cloud infrastructure (GCP). Experience Requirements...
BayOne Solutions
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
$117k - $209.33k
...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure... ...and security requirements ~ Experience with containers, Kubernetes, cloud-native architectures, APIs, load balancing,...
For contractors
Autodesk
San Francisco, CA
3 days ago
Solutions Engineer: Kubernetes, Python & Deployments
A tech company specializing in document intelligence is seeking a Solutions Engineer to work in San Francisco. This role involves deploying and operating their services in Kubernetes environments, collaborating directly with customer teams, and optimizing extraction pipelines...
Trypulse
San Francisco, CA
2 days ago
Senior Site Reliability Engineer
...at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software... ...-end project delivery Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience...
Immediate start
Remote work
Worldwide
OutSystems
San Francisco, CA
4 days ago
Site Reliability Engineer
$125k - $165k
...Site Reliability Engineer TELCOR Inc, a leading innovator in laboratory software, is looking for a Site Reliability Engineer to join our TELCOR... ...with queuing systems ~2+ years of experience with Kubernetes ~ Experience with AWS ~2+ years of experience with Terraform...
Work at office
Remote work
TELCOR
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability... ...with containerization (Docker), and orchestration (Kubernetes) ~ Strong knowledge of Linux systems, networking, and systems...
Alembic Technologies
San Francisco, CA
2 days ago
Senior Site Reliability Engineer
$181.69k - $213.75k
...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects... .... Our stack is Python, Java, Terraform, gRPC, Docker, Kubernetes, Postgres, running on AWS. Come join us! Cloud...
Full time
Work at office
Carta
San Francisco, CA
2 days ago
Site Reliability Engineer
$260k - $300k
...makers of Devin, the first AI software engineer, and Windsurf, the AI-native IDE. Together... .... You will own both the production reliability of our user-facing products and the... ...GCP, or Azure), container orchestration (Kubernetes), and infrastructure as code (Terraform...
Cognition Corp
San Francisco, CA
2 days ago
Site Reliability Engineer
$230k - $310k
...daily users while enabling our engineering teams to ship fast. You'll... ...and tooling that improves reliability and partnering with engineering... ...'ll Bring ~5+ years in site reliability engineering,... ..., containerization (Docker, Kubernetes), and database performance at...
Full time
Work at office
Work from home
Gamma
San Francisco, CA
1 day ago
Senior Site Reliability Engineer
$159.2k - $301.6k
...Graphs on the cloud. In this reliability-focused role, you will own... ...with cluster orchestrators like Kubernetes and building on top of AWS... ...'ll partner with the backend engineers building these APIs to make sure... ...~5-10 years of experience in site reliability engineering,...
Temporary work
Local area
Worldwide
Adobe
San Francisco, CA
2 days ago
Site Reliability Engineer
...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI... ...operational tooling. Familiarity with containerization (Docker, Kubernetes a plus). Embedded systems experience — reading firmware...
Remote work
Specter Services LLC
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
$166.9k - $225.9k
...operates as both a central engineering function and an embedded reliability practice. You'll be part... ...engineering leads and staff engineers to define SLOs... ...years of experience in Site Reliability Engineering,... ...AWS ECS Fargate and/or Kubernetes ~ Experience working with...
Work at office
Immediate start
Worldwide
Monday to Friday
Flexible hours
Drata Inc
San Francisco, CA
4 days ago
Site Reliability Engineer (SRE)
...management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring... ...with containerization and orchestration (Docker, Kubernetes, Helm). Strong Linux system administration and networking...
Temporary work
Worldwide
Air Apps
San Francisco, CA
4 days ago
Site Reliability Engineer
...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the... ...infrastructure for our users that scales, is reliable, and makes the complexities of operating... ...— you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like...
Permanent employment
Shift work
Arena AI
San Francisco, CA
8 hours ago
Site Reliability Engineer
...responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth... ...blog. As an SRE, you will Scale our Kubernetes-based control plane that processes billions of events...
Worldwide
Home office
Flexible hours
Superhuman
San Francisco, CA
1 day ago
Kubernetes Platform Engineer for Scalable Systems
A scaling startup is hiring an Infrastructure Engineer to manage a Kubernetes-based platform. You will be responsible for improving developer velocity, production reliability, and cloud infrastructure using Infrastructure-as-Code. Key tasks include evolving the Kubernetes...
Harrison Clarke
San Francisco, CA
3 days ago
Senior Staff Site Reliability Engineer
$220k - $235k
...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move... ...+ years of coding experience ~ Expert knowledge of Kubernetes and Google Cloud Platform (or similar provider) ~ Ability...
Full time
Contract work
Work at office
Ironclad Inc
San Francisco, CA
4 days ago
Staff Site Reliability Engineer
$200k - $260k
...Infrastructure Team as a technical leader driving reliability, automation, and scalability across the... ...practices across teams, mentor senior engineers, and be a primary escalation point for... ...you do ~10+ years of experience with Kubernetes/Docker in at least one top-tier cloud...
Casual work
Work at office
Remote work
Flexible hours
Sight Machine
San Francisco, CA
1 day ago
Manager, Site Reliability Engineering
$204k - $281k
...in on this mission. If you are too, let’s talk. Manager, Site Reliability Engineering San Francisco, California Okta authenticates, authorizes... ...expertise in cloud‑native architectures, containerization (Kubernetes), IaC (Terraform), and CI/CD pipelines. Strong background...
Permanent employment
Worldwide
Flexible hours
Okta, Inc.
San Francisco, CA
1 day ago
Lead Site Reliability Engineer
...Lead Site Reliability Engineer Stuut is transforming accounts receivable for B2B companies—making collections smarter and faster for companies... ...resilient, scalable cloud infrastructure across AWS and Kubernetes, ensuring systems are secure, fault-tolerant, and cost-...
Full time
Flexible hours
Stuut
San Francisco, CA
4 days ago
Staff Site Reliability Engineer
...Job: Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary... ..., compute, and storage. ~ Deep expertise in Kubernetes and the cloud-native ecosystem. ~ Fluency in at least...
United IT Solutions
San Francisco, CA
8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer - Kubernetes. Be the first to apply!