Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

SRE Lead

ReqRoute,Inc

Job Title: SRE Lead

Location: Woonsocket, RI - hybrid schedule

Job Description / Responsibilities

  • 8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility
  • Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents - structured leadership updates, not just participant involvement
  • Experience tuning and validating time-series anomaly detection models in a production observability context - this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role
  • Strong programming proficiency in Python, React, and Java at production quality - capable of writing operational tooling that other engineers will rely on
  • Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services
  • Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least two of the log aggregation solution (Loki, Splunk, Elasticsearch)
  • Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments
  • Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.
  • Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development.
  • Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipeline : Apache Airflow and Tidal.

Preferred Qualifications

  • Experience owning Production Readiness Reviews or service launch gates.
  • Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics, and reporting frameworks: Google BigQuery, PostgreSQL.
  • Hands-on chaos or fault injection experience.
  • TIC (Technical Incident Commander) certification or equivalent structured incident command training
  • Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact
  • LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) - design or implementation experience
  • Experience with streaming data platforms: Kafka.
  • Experience with service mesh and traffic management: Istio, Envoy.
  • Infrastructure-as-code proficiency at production scale: Terraform or Ansible
Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to SRE Lead. Be the first to apply!