Senior Site Reliability Engineer (US)
$130k - $170kClimavision
Senior Site Reliability Engineer
Remote | US
About Climavision
At Climavision, we're rebuilding climate technology from the ground up and changing the way we see weather. We merge the power of our proprietary, high-resolution weather radar and satellite network with advanced weather prediction modelling and decades of industry expertise to reduce existing coverage gaps and drastically improve forecasting ability. Our revolutionary new approach to climate technology weather solutions is poised to help reduce the economic risks of climate change on companies, governments, and societies alike. We are backed by The Rise Fund, the world's largest global impact platform committed to achieving measurable, positive social and environmental outcomes alongside competitive financial returns. Climavision is headquartered in Louisville, KY, with research and development operations in Raleigh, NC.
The Work
Are you an experienced Site Reliability Engineer who thrives at the intersection of software engineering and production operations? Do you take pride in keeping mission-critical customer systems reliable under real-world operational pressure? Are you looking for an opportunity to own production reliability for a modern hybrid infrastructure platform spanning cloud, colocation, and edge environments?
If so, we have an exceptional opportunity for you.
Climavision is seeking a Senior Site Reliability Engineer to contribute towards reliability, operational excellence, and production resilience across the company's platform and data services. This role sits on a shared SRE team that supports the full business rather than a single product line, covering both the radar network and the weather intelligence sides of the company as priorities shift. A central focus of this role is building the observability layer that puts the health of the full fleet in one place, and then automating recovery so that systems heal themselves. Multi-cluster and multi-replica high availability across our distributed edge fleet remains a core part of the work.
This is a hands-on engineering role for someone who is equally comfortable troubleshooting Kubernetes clusters, leading incident response, and improving operational maturity across the organization. The successful candidate will combine deep production operations expertise with a disciplined approach to reliability engineering and strong automation skills.
Climavision operates a hybrid infrastructure footprint spanning Microsoft Azure, colocation data centers, and edge Kubernetes clusters, deployed alongside weather radar systems. This role will drive production reliability across Azure, colocation, and edge environments. Right-sizing cluster resources and migrating workloads off Azure to reduce spend are active priorities for the team.
35% Kubernetes Platform Reliability and Operations
30% Production Reliability Engineering and Incident Response
20% Observability, Monitoring, and Alerting
15% Automation, Recovery, and Cost Optimization
Primary Responsibilities:
• Own production reliability for Climavision's customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
• Work as part of a shared SRE function supporting the whole company rather than a single product line, taking on work across both the radar network and the weather intelligence sides of the business as priorities shift.
• Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability.
• Build and own the observability layer for the fleet. Today the underlying metrics exist but are only reachable from the command line inside each cluster. This role is responsible for surfacing that data in shared dashboards and building the alerting that tells the team something is going wrong before a customer does.
• Design and build automated recovery and self-healing for production systems, so that common failure modes are detected and remediated without human intervention.
• Optimize cluster resourcing and cost, including right-sizing workloads and nodes and supporting the migration of workloads off Azure to reduce spend.
• Support and coordinate production incident response efforts, including troubleshooting, mitigation, communication, and postmortem analysis.
• Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems.
• Drive multi-replica and multi-cluster high availability across Climavision's services, including workload placement, scheduling, and deployment patterns that allow services to run safely as multiple replicas across multiple clusters.
• Contribute to the multi-cluster high-availability strategy across Climavision's hybrid fleet, including active-active and active-passive failover behavior, traffic routing, data replication considerations, and graceful degradation when a cluster becomes unavailable.
• Operate and improve Climavision's self-managed Kubernetes platform spanning cloud-hosted, colocation, and edge clusters, with a focus on availability, resiliency, recovery, and operational performance.
• Ensure Kubernetes platform lifecycle activities including upgrades, patching, cluster health, node management, and production change management are executed in a manner that preserves service availability and minimizes customer-facing risk.
• Improve reliability and operational maturity of production platform services, including observability, autoscaling, ingress, and distributed storage. Partner with the teams responsible for the underlying networking and security primitives rather than owning those areas directly.
• Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency, including autoscaling behavior, workload placement, resource management, and graceful degradation strategies.
• Partner with software engineering teams across the company to improve production readiness, resiliency patterns, deployment safety, and operational visibility before services reach production.
• Maintain and improve deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation supporting safe and repeatable production releases.
• Support and evolve Climavision's observability platform, including metrics, logging, distributed tracing, dashboarding, and alerting.
• Conduct performance engineering and capacity-planning efforts for customer-facing services during peak weather-event demand.
• Help facilitate blameless postmortem reviews and drive operational follow-up items through completion.
• Improve disaster recovery, failover, and business continuity capabilities across cloud, colocation, and edge environments.
• Drive operational excellence initiatives, including automation, reduction of operational toil, game days, production readiness reviews, and reliability best practices.
• Contribute as a senior technical resource and mentor on reliability engineering and production operations practices.
On-Call Expectation:
Climavision operates customer-facing production systems under contractual SLAs that do not pause outside business hours. The Senior Site Reliability Engineer will participate in a rotating on-call schedule made up of two separate rotations:
• A weekday rotation. The engineer on a weekday shift is the first point of contact for production incidents and for engineering teams needing support during the business week.
• A separate weekend rotation, so that the engineer carrying weekday support is not also carrying the weekend.
At current and planned team size, engineers can expect a weekday shift roughly every five weeks and a weekend shift roughly every five weeks. The two are scheduled as far apart from each other as the rotation allows, so that a weekday shift and a weekend shift do not fall close together.
Qualifications
• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered.
• Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability.
• Deep, hands-on experience operating native Kubernetes. Managed distributions such as AKS and EKS are acceptable, but experience running native or self-managed Kubernetes is strongly preferred and is the primary technical requirement for this role.
• Demonstrated experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost without sacrificing reliability. Be prepared to walk through a specific cluster optimization project you led.
• Demonstrated experience increasing operational visibility, including building dashboards, metrics pipelines, and alerting in an environment where little or none existed before.
• Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling considerations.
• Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment.
• Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities.
• Experience diagnosing and resolving production incidents across application, platform and Kubernetes infrastructure layers, including workload scheduling, storage, ingress, and cluster-level failures.
• Experience operating Kubernetes outside of strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure.
• Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems.
• Strong understanding of infrastructure automation and Infrastructure as Code concepts using tools such as Terraform and Ansible.
• Experience supporting CI/CD and production deployment pipelines. GitHub Actions is used for CI/CD at Climavision.
• Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies.
• Experience operating distributed systems and microservice-based architectures in production environments.
• Working knowledge of Microsoft Azure infrastructure.
• Strong troubleshooting skills across infrastructure, application, and platform layers.
• Demonstrated experience participating in a structured production on-call rotation supporting business-critical systems.
• Working familiarity with Jira, Confluence, and Microsoft Entra, which the team uses day to day for ticketing, documentation, and authentication.
• Strong written and verbal communication skills, including incident documentation and postmortem authoring.
• Experience working in start-up, scale-up, or other fast-moving engineering environments, and comfort with the pace and ambiguity that comes with them.
Nice to have, but not required:
• Experience operating Kubernetes platforms using RKE2 and Rancher, which is how Climavision manages its clusters.
• Experience with Octopus Deploy.
• Experience with SOC 2 or comparable security auditing and compliance work.
• Experience supporting hybrid cloud and colocation infrastructure environments.
• Experience with service mesh technologies such as Istio.
• Experience with Kubernetes-native storage platforms such as Longhorn.
• Experience operating PostgreSQL or PostGIS in Kubernetes environments.
• Experience with distributed messaging systems such as RabbitMQ or NATS.
• Experience supporting GPU-enabled workloads in Kubernetes.
• Familiarity with reliability engineering practices, including SLIs, SLOs, error budgets, and operational maturity metrics.
Physical Demands & Work Environment:
• This is a full-time, exempt position
• Fully Remote - United States
• This job requires frequent use of a computer to complete tasks, attend meetings, and communicate via Microsoft Teams.
Once you land this position, you'll get to enjoy:
• Benefits of a dynamic and growing organization
• A challenging, hands-on role that will have real impact on the business
• Competitive compensation
• Comprehensive benefits package
• 401(k) Savings Plan
• Medical/Dental/Vision Benefits
• Health Savings Account (HSA) and Flexible Spending Account (FSA)
• Unlimited Paid Time-off
• 11 Paid Holidays
• Paid Parental Leave
• Company Paid Short-term Disability (STD)
• Company Paid Long-term Disability (LTD)
• Company Paid Life Insurance
The salary range for this position is $130,000-170,000 annually, however Climavision considers several factors when extending an offer of employment including but not limited to, the applicant's education, experience, the responsibilities of the role, training, knowledge, skills, and abilities, as well as internal equity and alignment with market data. Any offer of employment is contingent on completion of a background check to company standard. Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without notice.
Climavision is an equal opportunity employer. All aspects of employment including the decision to hire, promote, discipline, or discharge, will be based on merit, competence, performance, and business needs. We do not discriminate on the basis of race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical condition, pregnancy, genetic information, gender, sexual orientation, gender identity or expression, veteran status, or any other status protected under federal, state, or local law.
$104k - $143k
...Role at Baxter Baxter is seeking a highly experienced Senior Site Reliability Engineer to join our dynamic team on a mission to save and sustain... ...sponsorship of an employment visa at this time. #LI-OM1 US Benefits at Baxter (except for Puerto Rico) This is...SeniorTemporary workWork experience placementRemote workWork visaFlexible hoursWeekend work$137.1k - $243.6k
...customers, and the world around us. As a Fortune 500 company and... ...on the Analytics Delivery Engineering team. As an SRE, you will help... ...in building a clean, scalable, reliable, and automated services framework... .... Please be aware of sites that may ask for you to input...SeniorFull timeWork experience placementWork at officeRemote workHome officeFlexible hours$174k - $252k
...pushing for changes that improve reliability and velocity.Practice... ...degree in Computer Science, Engineering, a related field, or equivalent... ...Computer Science or Engineering.Site Reliability Engineering (SRE)... ...relevant education or training. US: $174000 - $252000 (USD) + 15...Senior- ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS, Kubernetes, Terraform, CloudFormation, Ansible, CloudWatch, Grafana, Datadog, OpenSearch, PagerDuty Requirements: Senior SRE...SeniorFull timeRemote work
- ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis...SeniorFull timeRemote work
$95.3k - $158.8k
Senior Site Reliability Engineer Are you passionate about building resilient, scalable systems that power mission-critical applications?Do you thrive... ...need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or...SeniorFull timeLocal areaRemote workWork from home- ...grow, and make a difference. Why Join Us FedRAMP now defines compliance as a set... ...job. Coalfire organizes its delivery engineering into capability-focused teams, and the... ...long after the build team leaves. As a Senior Site Reliability Engineer you own one operational...SeniorFull time
- ...Senior Sre We're hiring a Senior SRE based in Latin America to work alongside our US-based engineering team, building out observability, on-call coverage, and deployment automation for a client with strict compliance requirements. We're specifically looking for someone...SeniorFull timeRemote work
- Senior Cloud Engineer Long term contract- 2+ years 100% remote in the continental US Our client, a premier national healthcare provider, is currently looking for a Senior Site Reliability Engineer on their Cloud Infrastructure team. You will be helping to maintain...SeniorLong term contractRemote work
$262k - $364k
...from SRE side, ensuring it is reliable, scalable, cost effective and... ..., while working closely with senior technical leads in the... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE)... ...relevant education or training. US: $262000 - $364000 (USD) + 25...Senior- ...Senior Site Reliability Engineer Company: ZetaChain Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Go, Python, Bash, Terraform, Ansible, Kubernetes, Docker, Linux, Prometheus, Grafana, Datadog, Loki, incident.io, AWS, GCP, Bare...SeniorFull timeRemote work
- ...Senior Site Reliability Engineer At Laravel, we don't just build tools; we build the foundation that empowers millions of developers to ship their... ...are looking for a Senior Site Reliability Engineer to help us scale that mission by ensuring our global infrastructure...SeniorRemote work
$150k - $170k
...Senior Site Reliability Engineer – Zip CoJoin to apply for the Senior Site Reliability Engineer role at Zip CoAt Zip, we build cloud-native software... ...engineering team.We offer a remote-first opportunity for US-based employees with the option to work in-person out of our...SeniorCasual workWork at officeRemote work$118k - $177k
...Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely . Everforth ECS is seeking talented professionals to join our... ...Description of Benefits Qualifications ~ Must be a US citizen with the ability to obtain Public Trust Suitability...SeniorRemote work$55 - $60 per hour
...#LI-Remote, prefer PST hours US, CA Pay Range: $55-60 Per... ...#LP Job Summary: This senior-level role focuses on building... ...candidate will possess strong cloud engineering judgment and the ability to... ...~ Strong understanding of reliability, security, cost management,...SeniorHourly payContract workRemote work$104.9k - $174.7k
...at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the... ...need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Formor please...SeniorFull timeWork at officeLocal areaRemote workWork from home$141.8k - $195k
...Why You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all... ...the Cribl herd? Learn more about the smartest, funniest, most passionate goats you’ll ever meet at cribl.io/about-us....SeniorTemporary workRemote work$150k - $220k
...Senior Site Reliability EngineerJob detailsDepartment / EngineeringRemoteFull-time$150,000 USD - $220... ...About The RoleAs a Senior Site Reliability Engineer, you own significant pieces of our... ...Type: Full Time Location: Fully Remote (US)## Benefits* Health / Dental / Vision insurance...SeniorFull timeRemote work- ...organizations maintain accurate, compliant, and reliable provider networks at scale. Our... ...the Role We're looking for a Senior Site Reliability Engineer who takes ownership seriously -... ...work. Your well-being matters to us. We provide 100% coverage of health,...SeniorRemote work
- ...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure... ...efficient, data-driven solutions. As a Senior Site Reliability Engineer, you will be responsible for:... ...grade solutions effectively. About us At Akamai, we make life better for...SeniorWork at officeRemote work
- ...Senior SRE Simple Life is the #1 AI-powered health coaching app for adults who want to... ...and the internal tooling that the rest of engineering relies on. Push the pace of innovation... ...build a future of a healthier world with us! About the Role: This is an...SeniorWork at officeRemote workFlexible hours
- ...barriers to application creation. About the role: Join our Site Reliability Engineering team and help ensure the reliability, scalability, and... ...Competitive Salary & Equity 401(k) Program with a 4% match ( US Only ) ⚕️ Health, Dental, Vision and Life Insurance...SeniorFull timeTemporary workWork at officeRemote workWorldwideFlexible hours
$180k - $230k
...About Us GridCARE is a leading venture-backed startup solving the most... ...Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our... ...work closely with platform and data engineering to keep high-throughput, data-intensive...SeniorWork at officeLocal areaImmediate startRemote work3 days per week- ...AI Senior SRE MOZN is a leading Enterprise AI company enabling... ...best work of your career, join us in shaping the future of... ...This role exists because most reliability toil (triage, root-causing, remediation... ...positive/negative rates, and engineer-hours of toil removed....SeniorRemote work
$140k - $180k
...About Us UJET leads the way in AI-powered contact center innovation, delivering a future-proof, cloud platform that... ...Learn more at Opportunity We’re looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. You’ll...SeniorWork experience placementLocal areaRemote workVisa sponsorshipWork visa- ...About The Role: We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure behind our multi-cloud SaaS platform. Most of our footprint runs on Microsoft Azure, built from the ground up around cloud architecture principles:...SeniorRemote workFlexible hours
- ...you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and...SeniorWork experience placementRemote workFlexible hours
- ...Senior Site Reliability Engineer Ciklum is looking for a Senior Site Reliability Engineer to join our team full-time in Ukraine. We are a custom... ...health support, and financial & legal consultations About us: At Ciklum, we are always exploring innovations,...SeniorFull timeWork at officeRemote work
- ...applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States. Join a globally distributed engineering... ...are ultimately made by humans. If you would like more information about how your data is processed, please contact us.SeniorRemote work
- ...Overview About the role You are a Senior Site Reliability Engineer based our Alpharetta, GA office. You lead Cloud Operations for Business Central... ...our diverse experiences, opinions and beliefs allows us to embrace what makes us unique and to use this as an asset...SeniorCasual workWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer (US). Be the first to apply!
- site reliability engineer Remote
- site reliability engineer remote Remote
- site reliability engineer sre Remote
- senior living director Remote
- senior manager customer operations Remote
- senior support engineer Remote
- senior product manager mobile Remote
- senior java developer Remote
- senior development engineer Remote
- senior software engineer ruby on rails Remote



