Lead Site Reliability Engineer
EPAM Systems, Inc.
We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript. Req.# Responsibilities Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, Google Cloud Platform Pub/Sub sinks, and OpenTelemetry (OTel) AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing Requirements AWS CDK & TypeScript for infrastructure provisioning and automation Agent & Cluster Agent deployment/management Monitors, Dashboards, Synthetics, and SLOs managed via Terraform Advanced constructs: Composite monitors, cardinality management, and retention controls OP2 / Vector (Observability Pipelines Worker) Log routing, enrichment, and redaction Splunk & Google Cloud Platform Pub/Sub sinks; OpenTelemetry standards CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice) AWS Lambda runtime monitoring & performance tuning PagerDuty integration and alert-to-page wiring from Datadog
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient... ...to resolution with very infrequent after-hours support. Lead blameless postmortems and implement long-term improvements...SuggestedFull time- ...EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this... ...SLOs and error budgets. • Modernization & Migration: Lead the technical execution of re-architecting and redeploying...SuggestedPermanent employmentFull timeH1bLocal areaRemote workShift work
$260k - $300k
...makers of Devin, the first AI software engineer. Our team is extremely talent-dense.... ...expects. You will own both the production reliability of our user-facing products and the... ...times. Incident Response and On-Call: Lead incident response with speed and clarity...Suggested$150k - $170k
...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud... ...0 days PTO every year ~ Generous paid parental leave ~ Leading family support policies ~ Company‑sponsored 401k match...SuggestedCasual workWork at officeRemote workFlexible hours$120k - $150k
...Site Reliability Engineer At Piper Sandler, we connect capital with opportunity to build a better future. We believe that diverse teams with... ...with cloud security posture or compliance concepts. As a leading investment bank, we enable growth and success for our...Suggested- ...Applications Deployment Responsible for reliability and support of Container Platform on-... ...Perform blameless RCA, partner with engineering and operation teams across the... ...Additional Skills : Automation Process Engineer,Site Reliability Engineer,Full Stack DeveloperThis...
$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...Local area$104k - $178k
...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The... ...incident reviews to minimize downtime and prevent recurrence. Leading technical projects from planning through deployment,...- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
- ...Site Reliability Engineer Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare...Relocation package
- ...SRE Engineer Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well) This is an SRE role supporting...
$182.8k - $247.3k
...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...Work experience placement- ...We are seeking a highly motivated Site Reliability Engineer (SRE) to join the Equity Trading Platform Engineering team, supporting critical trading... ...tuning, capacity planning, and fault isolation. Lead or participate in incident response, root-cause analysis, post...Permanent employmentWork at officeAfternoon shift
$189k - $283.6k
...proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...0) services. In this role, you will lead incident command, coordinate mitigation,... ...strong desire to perform and grow as an engineer ~5+ years of software development...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work$120k - $180k
...people, and works with high-profile manufacturers including leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield opportunity to architect the path from AWS to on-premises...Permanent employmentFull timeRelocation package- ...Software Reliability Engineer Good software has to run where customers need it. For many of Retool's largest customers, that means running... ...products, especially when they introduce new dependencies Lead through ambiguity, make careful risk calls, and communicate clearly...
$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...productivity and personalized customer experiences. Adobe's industry-leading offerings including Adobe Acrobat Studio, Adobe Express,...Temporary workLocal areaRemote workWorldwide- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...CD, ArgoCD). Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis....Flexible hours
- ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure...
- ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant... ...in and improve on-call rotations and incident response. Lead incident triage, mitigation, and resolution in real time....Full timeWork at officeRemote workFlexible hours2 days per week
- ...Chariot Engineering Hire Chariot's engineering hire will be responsible for taking the Chariot platform to the next level. You will lead the evolution of our banking product, DAF product and technology strategy, along with building a growing team of software engineers...Work experience placementWork at officeWork from homeMonday to FridayMonday to Thursday
- ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications...Flexible hours
- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....Local area
$130k - $200k
...who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll...Shift work$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...Work experience placementWork at officeLocal areaRelocation3 days per week$200k - $240k
...better patient care. Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health,... ...to meet you! About the role We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their...Work at office3 days per week- ...Site Reliability Engineer Our Client, a multinational telecommunications technology company is seeking a Site Reliability Engineer (SRE I) to join our Video Platform Engineering Team. As a Level 1 SRE, you will work closely with senior engineers to respond to incidents...Temporary work
- ...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank... ...to reduce manual steps and speed repeatable triage. Lead L1/L2 production support using SRE practices: quickly triage...Shift work
$80k - $95k
...applications, and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak... .... At RANE, we make it possible to grow, contribute, and lead—while living a life that works for you. We are an equal...Remote workVisa sponsorshipWork visa$191k - $226k
...anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the... .... You will run the machine: defining and upholding SLOs, leading incident response, and driving the automation and standards...Remote workWork visaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead infrastructure engineer New York, NY
- lead security engineer New York, NY
- lead engineer New York, NY
- lead operating engineer New York, NY
- lead system engineer New York, NY
- lead algorithm engineer New York, NY
- lead web developer New York, NY
- lead product engineer New York, NY
- lead network engineer New York, NY
- lead app. developer New York, NY



