Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Mirantis

Senior Site Reliability Engineer

We are looking for a senior Kubernetes-focused DevOps/SRE engineer to own both the developer platform and a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run the environments, pipelines, and infrastructure tooling that our engineering teams across the US, Europe, and APAC depend on to ship daily.

This role spans both sides of the line. You will make our development, test, and pre-production clusters fast and reproducible, harden the Helm and CI/CD path from commit to release, and carry operational ownership — including on-call — for one of our smaller customer-facing production regions, under real availability commitments. That production experience makes you the internal expert on how k0rdent AI is deployed and operated — the person other teams consult, including the teams running our larger regions. Working within an agile framework, you will directly shape how quickly and safely changes reach production, and be accountable for how they behave once there.

Main Responsibilities:

  • Own the Kubernetes footprint across development, CI, pre-production, and one customer-facing production region — local kind clusters, shared dev and QA environments, and multi-cluster/multi-region topologies.
  • Operate your production region against defined SLOs: capacity and upgrade planning, patching, backup and restore, disaster recovery drills, and participation in an on-call rotation.
  • Lead incident response for your region — detection, mitigation, customer-impact assessment, root-cause analysis, and blameless postmortems that feed fixes back into the platform.
  • Build and maintain Helm charts and umbrella releases for the platform's services and dependencies, including versioning, values hygiene, and upgrade paths.
  • Own the CI/CD pipelines end to end — build, test, image publishing, chart packaging, release cutting, and hotfix/backport flows.
  • Automate environment bootstrap and seeding so any engineer can bring up a full stack — control plane, identity, gateway, database, workflow engine — with one command.
  • Operate and troubleshoot the supporting stack across test and production: PostgreSQL, Temporal, Keycloak, API gateway, message broker, and observability components.
  • Build observability and diagnostics — metrics, dashboards, alerting, log and audit access — that serve both engineering environments and production operations.
  • Consult with product teams and with the teams operating our larger regions on deployment topology, GPU and resource scheduling, RBAC, networking, and failure modes; validate upgrade and migration procedures and hand over runbooks.
  • Enforce security and tenant isolation in production: least-privilege access, secret handling, certificate and TLS lifecycle, image and dependency scanning, and audit evidence for compliance reviews.
  • Drive infrastructure as code and repeatability — no snowflake environments, no undocumented manual steps.
  • Mentor engineers on Kubernetes and operational practice, and raise the team's bar through review and documentation.
Qualifications

Required Skills/Abilities:

  • 10+ years in DevOps, SRE, platform, or infrastructure engineering, including production ownership of customer-facing Kubernetes environments.
  • Expert-level Kubernetes: workloads, networking, storage, RBAC, resource management, CRDs and operators, and cluster upgrades — able to debug from kubectl and cluster internals rather than dashboards alone.
  • Proven incident response under SLA pressure — on-call rotations, escalation paths, postmortems, and follow-through on corrective action.
  • Strong CI/CD engineering — pipelines as code, reproducible builds, artifact and release management (GitHub Actions or equivalent).
  • Solid scripting and automation ability, and enough Go familiarity to read service code, trace a failure into it, and file a precise bug.
  • Experience running the stateful supporting stack — relational databases, identity providers, gateways, and message brokers — in Kubernetes, including backup, restore, and upgrade.
  • Track record as a technical consultant to other engineering teams: clear runbooks, design feedback, and incident write-ups across global time zones (strong written English).

Must Have:

  • Kubernetes Native: Kubernetes at scale, Cluster API, controllers/operators, Docker, and Helm chart authoring and lifecycle management.
  • Delivery: GitHub Actions or equivalent CI/CD, container registries, versioned release and backport workflows.
  • Infrastructure as Code: Terraform, Ansible, or equivalent, plus GitOps tooling (Argo CD, Flux).
  • Identity & API Management: Keycloak and API gateway operation — routing, plugins, TLS, rate limiting.
  • Data & Messaging: PostgreSQL operations and migrations, plus streaming/message-broker platforms (Kafka or equivalent).
  • Observability: Prometheus, Grafana, centralized logging, and alerting tied to SLOs.
  • Cloud: AWS— networking, IAM, load balancing, and managed Kubernetes

Nice to Have:

  • k0s or k0rdent ecosystem experience.
  • GPU infrastructure on Kubernetes — device plugins, node feature discovery, scheduling and sharing of accelerators.
  • Temporal operations — namespaces, workers, schema upgrades.
  • Bare-metal provisioning (Metal³ / BareMetalHost) or on-prem/OpenStack environments.
  • Multi-region topologies, service mesh, or cross-cluster networking.
  • Python for test harnesses and automation; experience with pytest-based E2E suites.
  • Load and performance testing of API platforms.
  • Policy enforcement (OPA/Kyverno), secret management, and pen-test remediation.
  • OpenTelemetry, distributed tracing, or formal SLO/error-budget practice.
  • Compliance exposure — SOC 2, ISO 27001, or similar audit support.
  • CNCF open-source contributions.

Education and Experience:

  • Bachelor's degree in Computer Science & Engineering or related field or 10 years related experience.
Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating... 
    Senior
    Full time

    Ridgeline

    New York, NY
    more than 2 months ago
  •  ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will... 
    Senior
    Remote work

    SRI Tech

    Plano, TX
    3 days ago
  • IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Senior
    Work at office
    Immediate start

    IXL Learning

    Raleigh, NC
    12 hours ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    3 days ago
  • $104.9k - $174.7k

    About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding cost optimization into infrastructure design, automation, monitoring, and operational processes. A FinOps SRE proactively identifies... 
    Senior
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Boca Raton, FL
    3 days ago
  • $182.8k - $247.3k

     ...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed... 
    Senior
    Work experience placement

    Duolingo

    Pittsburgh, PA
    1 day ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    12 hours ago
  • $152.6k - $191.5k

     ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...and continuous improvement.Position Summary: The Senior Site Reliability Engineer acts as an advanced senior individual... 
    Senior
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, MI
    3 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    New York, NY
    10 hours ago
  •  ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform, CloudFormation, Ansible, CloudWatch, Grafana, Datadog, OpenSearch, PagerDuty Requirements: Senior SRE... 
    Senior
    Full time
    Remote work

    CyberArk

    United States
    12 hours ago
  •  ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver... 
    Senior
    Full time
    Flexible hours

    Proofpoint

    Florida
    3 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Gainesville, FL
    4 days ago
  •  ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The...  ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and... 
    Senior

    Intercontinental Exchange

    Jacksonville, FL
    2 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    12 hours ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $15k

     ...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    1 day ago
  • $190.8k - $267.1k

     ...helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser...  ...team partners closely with Ads Engineering to improve reliability, scalability, operational...  ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,... 
    Senior
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    3 days ago
  • The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also... 
    Senior
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    4 days ago
  •  ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis... 
    Senior
    Full time
    Remote work

    Sphera

    United States
    12 hours ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    4 days ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Senior
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    4 days ago
  • $91.7k - $163.7k

     ...potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud environment in both the commercial and government cloud. The... 
    Senior
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Eden Prairie, MN
    4 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    12 hours ago
  • $127k - $249k

    The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB...  ...a pivotal role in engineering the reliable, globally connected, multi-cloud...  ...Role OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a strong... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    3 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    4 days ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the... 
    Senior
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    3 days ago
  • $80k - $140k

    Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and... 
    Senior
    Full time
    Flexible hours
    Shift work

    Royal Bank of Canada

    Minneapolis, MN
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!