Lead Site Reliability Engineer
EPAM Systems, Inc.
Skip To Main ContentBack to SearchRemote in United States of America: New YorkSite Reliability Engineering& 4 othersLooking for something else?Find a vacancy that works for you. Send us your CV to receive a personalized offer.Find me a jobWe are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.Req.#1080523852Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architecturesPlatform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScriptTelemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and FirehoseObservability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routingAWS CDK & TypeScript for infrastructure provisioning and automationAgent & Cluster Agent deployment/managementMonitors, Dashboards, Synthetics, and SLOs managed via TerraformAdvanced constructs: Composite monitors, cardinality management, and retention controlsOP2 / Vector (Observability Pipelines Worker)Log routing, enrichment, and redactionSplunk & GCP Pub/Sub sinks; OpenTelemetry standardsCloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)AWS Lambda runtime monitoring & performance tuningPagerDuty integration and alert-to-page wiring from Datadog
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient... ...to resolution with very infrequent after-hours support. Lead blameless postmortems and implement long-term improvements...SuggestedFull time- ...EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this... ...SLOs and error budgets. • Modernization & Migration: Lead the technical execution of re-architecting and redeploying...SuggestedPermanent employmentFull timeH1bLocal areaRemote workShift work
$141k - $216.6k
...building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational... ...of ambiguous technical problems with minimal direction.Leading operational improvements rather than simply executing assigned...SuggestedWork experience placementWork at office$158.5k - $172k
...velocity energy of a powerhouse startup.As a leading U.S. ordering and delivery marketplace,... ....About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and automation...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week$167.7k - $245.2k
...assurance insights within Cisco’s leading Networking, Security,... ...effective.We’re looking for talented engineers with a software or operations... ...teams to ensure the reliability, performance and security of... ...Please see the Cisco careers site to discover more benefits and...SuggestedFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the...Work at officeLocal areaImmediate start$45 - $85 per hour
DescriptionThe Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.We're looking for engineers to... ...law.About TEKsystems and TEKsystems Global Services We’re a leading provider of business and technology services. We...Contract workTemporary work- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank... ...tooling to reduce manual steps and speed repeatable triage.Lead L1/L2 production support using SRE practices: quickly...Shift work
$120k - $150k
...to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About... ...with cloud security posture or compliance concepts.As a leading investment bank, we enable growth and success for our clients...Full time$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express,...Full timeTemporary workLocal areaRemote workWorldwide$207k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives. Mentor other engineers and contribute to the engineering community through documentation...Full timeWork at office$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...Work experience placement$190k - $260k
Who are we?Cohere is the leading security-first enterprise AI company... ...is a team of researchers, engineers, designers, and more, who are... ...high-performance, scalable and reliable machine learning systems? Do... ...? We are looking for a Site Reliability Engineer to join...Full timeWork experience placementWork at officeLocal areaRemote workHome office$194k - $267k
...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is... ...technical challenges. You will serve as a key technical lead within the EPG SRE organization, partnering with software engineers...Local areaWorldwideFlexible hours$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$150k - $220k
...and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative... ...achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems...Local area$182k - $250.8k
...Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great... ...of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability...Permanent employmentLocal areaRemote workWorldwideFlexible hoursWeekend workWeekday work- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership...
$131k - $164k
Position OverviewWe are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation... ...they need to drive greater impact and accountability - to lead with purpose. Our employees are passionate, smart, and...Work at officeLocal areaVisa sponsorshipFlexible hours$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$195k - $275k
Morgan Stanley is a leading global financial services firm providing a wide range of investment banking, securities, investment management... ...& Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift,...Temporary workWork at officeWorldwideNight shift$130k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible... ...your journey here.Securities Frontline Site Reliability Engineers (SREs) play a critical role... ...serve to grow. Founded in 1869, we are a leading global investment banking, securities...Full timeTemporary workPart timeImmediate start- ...Job Title WHAT YOU'LL DO DAY-TO-DAY: The engineer will be responsible for developing and implementing new systems and services in the areas of infrastructure monitoring, configuration management, and automation. Additional responsibilities include upgrading and/or...
- ...Triomics Backend Engineer Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained staff manually reading unstructured patient records - pathology reports, clinical notes, genomic panels - to power workflows like...Day shift
- ...Site Reliability Engineer TXSE is building the next-generation exchange infrastructure to support transparent, efficient, and resilient capital markets. With SEC approval and $275MM in funding, we are currently hiring a Site Reliability Engineer to help with a greenfield...Currently hiring
$150k - $160k
Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join... ...fast, resilient, and optimized at the edge.Responsibilities:Lead development team initiatives to:Architect and maintain Cloudflare...Work at officeLocal area$111k - $218k
...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency...Local areaWorldwideFlexible hours$115k - $125k
...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help...Work experience placementWork at officeWork from homeMonday to FridayFlexible hoursShift workWeekend workAfternoon shiftEarly shift$105k - $300k
...Site Reliability Engineer At Citadel, a leading investor in the world's financial markets, we aim to win together as one team to earn the long-term trust of our capital partners and each other. Our collaborative approach allows technologists to grow alongside other...$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ..., while embedding directly with product engineering teams and their leadership to drive... ...comprises hundreds of libraries and services.Lead initiatives around reliability,...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead algorithm engineer New York, NY
- lead web developer New York, NY
- lead network engineer New York, NY
- lead infrastructure engineer New York, NY
- lead system engineer New York, NY
- lead operating engineer New York, NY
- lead engineer New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- site reliability engineer sre New York, NY


