Staff Site Reliability Engineer - Kubernetes
$194k - $267kOkta
Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence. This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.Workforce Identity Cloud
Okta Workforce Identity Cloud (WIC) provides easy, secure access for your workforce so you can focus on other strategic priorities—like reducing costs, and doing more for your customers.
If you like to be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from you. The ideal candidate is someone who exemplifies the ethics of, “If you have to do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.
Position Overview:
The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on architecting and managing reliable, scalable, and secure Kubernetes-based platforms on AWS, ensuring high availability and performance while optimizing costs and automation. The ideal candidate will have hands-on experience with AWS infrastructure, Kubernetes platform creation, Helm charts, Karpenter scaling, and Istio service mesh.
Key Responsibilities:
- Kubernetes Platform Creation: Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience and operational efficiency.
- AWS Infrastructure Management: Build, manage, and optimize AWS cloud infrastructure, including EKS,ECS, S3, VPCs, RDS, IAM, and more. Implement best practices for cost management, scaling, and security within AWS.
- Helm Management: Utilize Helm to automate and streamline the deployment of applications and services to Kubernetes clusters. Create, maintain, and manage Helm charts for production-ready deployments.
- Karpenter Implementation: Implement and manage Karpenter to dynamically scale Kubernetes clusters in response to workload demands.
- Istio Service Mesh Management: Configure and manage Istio to provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained traffic management, service discovery, and policy enforcement.
- Platform Automation & Scaling: Automate the deployment, scaling, and management of infrastructure and applications. Work with CI/CD pipelines to ensure a seamless flow from development to production with minimal downtime.
- Incident Management & Troubleshooting: Respond to incidents, troubleshoot, and resolve system issues related to performance, availability, and security in a timely and effective manner.
- Security & Compliance: Design and implement secure cloud infrastructure with appropriate access controls, network security, and compliance frameworks.
- Documentation & Knowledge Sharing: Create and maintain detailed documentation for Kubernetes platform setup, operational procedures, and best practices. Promote knowledge sharing across teams.
Required Qualifications:
- 4+ years of experience with Kubernetes/Helm;
- 4+ years of Experience with Terraform.
- 5+ years of Experience with AWS
- Experience with multi-region cloud environments.
- Proven experience with AWS (EC2, RDS, S3, CloudFormation, IAM, etc.) and solid understanding of cloud-native architectures.
- Strong expertise in Kubernetes platform creation, management, and optimisation (e.g., setting up highly available clusters, networking, and storage).
- Hands-on experience with Helm for Kubernetes application deployment and management.
- Practical experience with Karpenter for dynamic scaling of Kubernetes clusters and optimising resource usage.
Expertise in managing and securing Istio for service mesh, including traffic management, security, and observability features. - Proficiency in CI/CD pipelines and automation tools (e.g., Jenkins, GitLab, CircleCI, Terraform, Ansible, Spinnaker).
Strong scripting and automation skills in Python, Bash, or Go for infrastructure management and platform automation. - Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack.
Preferred Qualifications:
- Understanding of security best practices for cloud platforms and Kubernetes (e.g., role-based access control (RBAC), encryption, and compliance frameworks).
- Familiarity with Docker and containerization principles.
- Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent professional experience).
- Certifications (Preferred): CKA (Certified Kubernetes Administrator), CKAD (Certified Kubernetes Application Developer), or AWS Certified DevOps Engineer are highly desirable.
Additional requirements:
- This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.
- Requires in-person onboarding and travel to our San Francisco, CA HQ office or our Chicago office during the first week of employment.
#LI-Hybrid
#LI-LSS1
requisition ID- (P16373_3396241)
The annual base salary range for this position for candidates located in the San Francisco Bay area is between: $194,000—$267,000 USD Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: .
The Okta Experience
We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.
Okta is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, marital status, age, physical or mental disability, or status as a protected veteran. We also consider for employment qualified applicants with arrest and convictions records, consistent with applicable laws. If reasonable accommodation is needed to complete any part of the job application, interview process, or onboarding please use this Form to request an accommodation. Notice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. In accordance with NYC Local Law 144, if you are an applicant or employee residing in New York City, please click here to view our full NYC AEDT Notice. Okta is committed to complying with applicable data privacy and security laws and regulations. For more information, please see our Personnel and Job Candidate Privacy Notice at .- ...world's most complex and mission-critical systems. As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the... ...Deploy, manage, and scale containerized applications using Kubernetes (EKS) and ECS. Develop and maintain infrastructure as code...Suggested
$160k - $250k
...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS... ...Containerization - Docker Container Orchestrators - Mesosphere/Kubernetes Scripting Languages - Python/Ruby/Node/Bash CI/CD...Suggested- ...from Altimetrik office ) Pay Rate -65 C2C Position: Site Reliability Engineer (SRE) - Backend/Cloud Must-Have Experience: Familiarity... ...deploying and managing services with Helm in Kubernetes environments . Strong background in Python and Object...SuggestedWork at officeRemote work
- ...in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the... ...such as Terraform Associate Certification and Certified Kubernetes Administrator Bonus Expertise in monitoring...Suggested
$127k - $249k
THE TEAM Platform Engineering is the department within SRE that is responsible for a range... ...these are our multi-cloud-provider Kubernetes infrastructure, networking, load balancing... ...components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...SuggestedWork at officeLocal areaRemote workWorldwideFlexible hours$142.3k - $263.3k
...Senior Site Reliability Engineer Apple Services Engineering Cloud Service Infrastructure team is one of the most exciting examples of Apple... ...variety of services based on open-source software, such as Kubernetes, Cassandra, Zookeeper, Kafka, Redis, etc, alongside...Relocation$100k - $170k
...real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads... ...with high-performance networking (InfiniBand, RDMA). • Kubernetes, plus virtualized or bare-metal environments. On-...Flexible hoursShift work- ...Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB Terraform...
- ...operate across the globe. We are looking for a Manager, Site Reliability Engineering to be part of revolutionizing these industries. What You... ...systematically eliminating manual operational work. Experience with Kubernetes container orchestration with production‑grade operational...
$153k - $179k
...developer velocity and increases system reliability by building the foundational platforms and tools that power Robinhood engineering. Within this group, the Compute team focuses... ...operating a highly available, scalable Kubernetes-powered container platform. We ensure that...Work at officeFlexible hoursShift work3 days per week$194k - $267k
...talk. We are seeking a highly technical Observability Site Reliability Engineer with a specialty in Google Cloud, to own and expand our Observability... .../IP, DNS, Load Balancing), and container orchestration (Kubernetes/GKE). Problem Solving: A data-driven approach to...Permanent employmentLocal areaWorldwideFlexible hours$194k - $267k
...are too, let's talk. The Team The Site Reliability team is dedicated to architecting and... ...that maximize platform reliability and engineering velocity. The ideal candidate is someone... ...Experience working with Docker and Kubernetes Experience with PKI / certificate...Local areaWorldwideFlexible hours$194k - $267k
...you are too, let's talk. The Team The Site Reliability team is dedicated to architecting and... ...that maximize platform reliability and engineering velocity. The ideal candidate is someone... .... Experience working with Docker and Kubernetes Experience with PKI / certificate...WorldwideFlexible hours$163.62k - $212.71k
...operating mainly in AWS with multiple Kubernetes clusters and thousands of servers. We... ...platforms, and processes that improve our engineering teams' productivity and streamline the... ...seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability,...Full timePart timeWork experience placementWork at officeLocal areaImmediate startRemote workWork from homeFlexible hoursShift work3 days per week1 day per week$116.36k - $155.15k
...and experience in system architecture and engineering disciplines. Specific technical... ...Cloud Platform. Support and troubleshoot Kubernetes clusters and containers. Support, troubleshoot... ...due diligence activities including site surveys, design, design review, bill of...Full timeTemporary workRemote work1 day per week- ...Software Engineer At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give our... ..., CSS SQL/NoSQL Database Devops - Gitlab CI, Docker, Kubernetes, AWS Strong knowledge of Agentic AI Strong knowledge of...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours
$139k - $229k
...Sr. Software Engineer, Systems Infrastructure Employment Type: Full-time, Hybrid Location... ...in technologies such as Hadoop, Spark, Kubernetes, Feather, GraphQL, gRPC, Apache Kafka, Pinot... ...this posting, you are leaving this site and going to a third‑party website where...Full timeFor contractorsWork experience placement- ...NoSQL (e.g. Cassandra etc.). Cloud Concepts. Frontend - Deep knowledge in React ReactJS concepts - hooks, redux, etc., MFA (Micro frontend Architecture), Webpack bundler, Babel transpiler. Good to have skills: Angular, Vue.JS, Cloud (AWS/Azure), Kubernetes, DevOps....
$88k - $136.9k
...through the most innovative, convenient, reliable, and secure payments network, enabling... ...Versatile, curious, and energetic Software Engineers who embrace solving complex challenges... ...new technologies such as Angular, React, Kubernetes, Docker, etc. Partnership: Experience...Work at officeLocal areaVisa sponsorshipRelocation package$157.6k - $197k
...Senior Software Engineer This role is office-based at our Bellevue, Washington office. Armada is the hyperscaler for the edge,... ...fabric that unifies a hybrid fleet of Azure Local, OpenShift, and Kubernetes clusters. You will also build the automation software that...Work at officeLocal areaFlexible hours- ...and Artifactory platforms Lead cloud-native transformations, Kubernetes migrations, and GitOps adoption Drive architecture for multi-... ...Design self-healing systems with automated remediation and chaos engineering practices Establish cloud-native best practices, including 12...
- ...Deep Sync is seeking a Senior Software Engineer to be a core member of our US engineering... ...platforms, ensuring they are scalable, reliable, and user-friendly. Collaborate with cross... ...technologies (e.g., Docker, Kubernetes ). Excellent problem-solving skills and...Temporary workWork at officeLocal areaRemote workFlexible hours
- ...Product Security Engineer At Snowflake, we are powering the era of the agentic enterprise... ..., building, testing, and maintaining reliable, and scalable software solutions. Experience... ...deploying and operating services on Kubernetes Experience with building production...
- ...What You’ll Do The Applications Engineering team in the Data Infrastructure organization... ...frontends to backend services running on Kubernetes— while integrating AI/LLM capabilities... ...gRPC) with a focus on performance and reliability Experience developing and deploying AI/...Permanent employmentTemporary workCasual workWork at officeFlexible hours
$52 - $85 per hour
...Responsibilities: Build safe, efficient, and reliable software systems which will include:... ...the latest developments in software engineering and related technologies Use tools... ...of cloud computing (containerization, Kubernetes, etc.) and cloud networking (virtual...Permanent employmentContract workRemote work$165k - $242k
...Senior Software Engineer, Developer Experience Livingston, NJ / New York, NY / Sunnyvale... ...shape developer velocity, system reliability, and the agentic developer experience across... ...environments. ~ Deep experience with Kubernetes. You understand how workloads are...Permanent employmentTemporary workCasual workWork at officeFlexible hours$117.2k - $313.7k
...Job Category Software Engineering Job Details About Salesforce... ...on our platform to be highly reliable, lightning fast, supremely secure... ...frameworks such as Kubernetes, Docker, Mesos Build and... ...have experience balancing live-site management, feature delivery,...$109k - $204k
...services, and systems that enable every engineer at CoreWeave to build and ship software... ...retrieval, and version management that reliably serve engineering teams at scale. Identify... ...of how they’re built. Familiarity with Kubernetes basics, such as how workloads are defined...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hours$68.91k - $161.54k
...Software Engineer Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you'd like... ...Familiarity with containerization technologies (e.g., Docker, Kubernetes). Experience with cloud platforms (e.g., AWS, Azure)....Permanent employmentFull timeContract workLocal area- ...you to apply it in production environments , using Spring Boot, cloud deployment (AWS, Azure), DevOps tooling (Jenkins, Docker, Kubernetes), and full- • stack development workflows. Hands-On Projects: You'll work on enterprise-level projects that simulate what...Full timeImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer - Kubernetes. Be the first to apply!
- assistant engineer Bellevue, WA
- senior staff systems engineer Bellevue, WA
- engineering aide Bellevue, WA
- staff engineer Bellevue, WA
- technology administrator Bellevue, WA
- site safety Bellevue, WA
- historic site Bellevue, WA
- IT site lead Bellevue, WA
- site leader Bellevue, WA
- site services specialist Bellevue, WA



