Senior Site Reliability Engineer
Mirantis
Senior Site Reliability Engineer
We are looking for a senior Kubernetes-focused DevOps/SRE engineer to own both the developer platform and a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run the environments, pipelines, and infrastructure tooling that our engineering teams across the US, Europe, and APAC depend on to ship daily.
This role spans both sides of the line. You will make our development, test, and pre-production clusters fast and reproducible, harden the Helm and CI/CD path from commit to release, and carry operational ownership — including on-call — for one of our smaller customer-facing production regions, under real availability commitments. That production experience makes you the internal expert on how k0rdent AI is deployed and operated — the person other teams consult, including the teams running our larger regions. Working within an agile framework, you will directly shape how quickly and safely changes reach production, and be accountable for how they behave once there.
Main Responsibilities:
- Own the Kubernetes footprint across development, CI, pre-production, and one customer-facing production region — local kind clusters, shared dev and QA environments, and multi-cluster/multi-region topologies.
- Operate your production region against defined SLOs: capacity and upgrade planning, patching, backup and restore, disaster recovery drills, and participation in an on-call rotation.
- Lead incident response for your region — detection, mitigation, customer-impact assessment, root-cause analysis, and blameless postmortems that feed fixes back into the platform.
- Build and maintain Helm charts and umbrella releases for the platform's services and dependencies, including versioning, values hygiene, and upgrade paths.
- Own the CI/CD pipelines end to end — build, test, image publishing, chart packaging, release cutting, and hotfix/backport flows.
- Automate environment bootstrap and seeding so any engineer can bring up a full stack — control plane, identity, gateway, database, workflow engine — with one command.
- Operate and troubleshoot the supporting stack across test and production: PostgreSQL, Temporal, Keycloak, API gateway, message broker, and observability components.
- Build observability and diagnostics — metrics, dashboards, alerting, log and audit access — that serve both engineering environments and production operations.
- Consult with product teams and with the teams operating our larger regions on deployment topology, GPU and resource scheduling, RBAC, networking, and failure modes; validate upgrade and migration procedures and hand over runbooks.
- Enforce security and tenant isolation in production: least-privilege access, secret handling, certificate and TLS lifecycle, image and dependency scanning, and audit evidence for compliance reviews.
- Drive infrastructure as code and repeatability — no snowflake environments, no undocumented manual steps.
- Mentor engineers on Kubernetes and operational practice, and raise the team's bar through review and documentation.
Qualifications
Required Skills/Abilities:
- 10+ years in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments.
- Expert-level Kubernetes: workloads, networking, storage, RBAC, resource management, CRDs and operators, and cluster upgrades — able to debug from kubectl and cluster internals rather than dashboards alone.
- Proven incident response under SLA pressure — on-call rotations, escalation paths, postmortems, and follow-through on corrective action.
- Strong CI/CD engineering — pipelines as code, reproducible builds, artifact and release management (GitHub Actions or equivalent).
- Solid scripting and automation ability, and enough Go familiarity to read service code, trace a failure into it, and file a precise bug.
- Experience running the stateful supporting stack — relational databases, identity providers, gateways, and message brokers — in Kubernetes, including backup, restore, and upgrade.
- Track record as a technical consultant to other engineering teams: clear runbooks, design feedback, and incident write-ups across global time zones (strong written English).
Must Have:
- Kubernetes Native: Kubernetes at scale, Cluster API, controllers/operators, Docker, and Helm chart authoring and lifecycle management.
- Delivery: GitHub Actions or equivalent CI/CD, container registries, versioned release and backport workflows.
- Infrastructure as Code: Terraform, Ansible, or equivalent, plus GitOps tooling (Argo CD, Flux).
- Identity & API Management: Keycloak and API gateway operation — routing, plugins, TLS, rate limiting.
- Data & Messaging: PostgreSQL operations and migrations, plus streaming/message-broker platforms (Kafka or equivalent).
- Observability: Prometheus, Grafana, centralized logging, and alerting tied to SLOs.
- Cloud: AWS— networking, IAM, load balancing, and managed Kubernetes
Nice to Have:
- k0s or k0rdent ecosystem experience.
- GPU infrastructure on Kubernetes — device plugins, node feature discovery, scheduling and sharing of accelerators.
- Temporal operations — namespaces, workers, schema upgrades.
- Bare-metal provisioning (Metal³ / BareMetalHost) or on-prem/OpenStack environments.
- Multi-region topologies, service mesh, or cross-cluster networking.
- Python for test harnesses and automation; experience with pytest-based E2E suites.
- Load and performance testing of API platforms.
- Policy enforcement (OPA/Kyverno), secret management, and pen-test remediation.
- OpenTelemetry, distributed tracing, or formal SLO/error-budget practice.
- Compliance exposure — SOC 2, ISO 27001, or similar audit support.
- CNCF open-source contributions.
Education and Experience:
- Bachelor's degree in Computer Science & Engineering or related field or 10 years related experience.
Additional Information
What does Mirantis offer you?
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
We are a Leader for Container Management in G2 (#2 after AWS)!
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating...SeniorFull time- ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will...SeniorRemote work
- IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SeniorWork at officeImmediate start
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
$104.9k - $174.7k
About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding cost optimization into infrastructure design, automation, monitoring, and operational processes. A FinOps SRE proactively identifies...SeniorFull timeLocal area$182.8k - $247.3k
...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
$152.6k - $191.5k
...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include... ...and continuous improvement.Position Summary: The Senior Site Reliability Engineer acts as an advanced senior individual...SeniorFull timeWork at officeDay shift$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...Senior- ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform, CloudFormation, Ansible, CloudWatch, Grafana, Datadog, OpenSearch, PagerDuty Requirements: Senior SRE...SeniorFull timeRemote work
- ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver...SeniorFull timeFlexible hours
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SeniorFull timeWork at officeLocal areaRemote workWork from home- ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The... ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and...Senior
$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...SeniorFull time$15k
...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorWork at officeLocal areaRemote work$190.8k - $267.1k
...helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser... ...team partners closely with Ads Engineering to improve reliability, scalability, operational... ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,...SeniorFor contractorsWork experience placement- The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also...SeniorFull timeWork experience placementRemote work
- ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis...SeniorFull timeRemote work
$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability...SeniorOngoing contractLocal area3 days per week$160k - $200k
...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident...SeniorTemporary workWork at officeLocal areaFlexible hours3 days per week- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SeniorTemporary workCasual workWorldwide
$91.7k - $163.7k
...potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud environment in both the commercial and government cloud. The...SeniorMinimum wageFull timeWork experience placementWork at officeLocal areaRemote work$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ...a pivotal role in engineering the reliable, globally connected, multi-cloud... ...Role OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a strong...SeniorLocal areaRemote workWorldwideFlexible hours$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...SeniorFull timeFor contractors- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the...SeniorFull timeWork at office
$80k - $140k
Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and...SeniorFull timeFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre United States
- site reliability engineer United States
- site reliability engineering manager United States
- site reliability engineer remote United States
- lead site reliability engineer United States
- senior facilities maintenance technician United States
- senior operations technician United States
- senior groundskeeper United States
- senior operations associate United States
- senior cloud service delivery manager United States


