Lead Site Reliability Engineer
$145k - $160kEPAM Systems, Inc.
We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.
Req.#1080523852
Responsibilities
-
Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures
-
Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript
-
Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose
-
Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)
-
AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing
Requirements
-
AWS CDK & TypeScript for infrastructure provisioning and automation
-
Agent & Cluster Agent deployment/management
-
Monitors, Dashboards, Synthetics, and SLOs managed via Terraform
-
Advanced constructs: Composite monitors, cardinality management, and retention controls
-
OP2 / Vector (Observability Pipelines Worker)
-
Log routing, enrichment, and redaction
-
Splunk & GCP Pub/Sub sinks; OpenTelemetry standards
-
CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)
-
AWS Lambda runtime monitoring & performance tuning
-
PagerDuty integration and alert-to-page wiring from Datadog
We offer
-
Medical, Dental and Vision Insurance (Subsidized)
-
Health Savings Account
-
Flexible Spending Accounts (Healthcare, Dependent Care, Commuter)
-
Short-Term and Long-Term Disability (Company Provided)
-
Life and AD&D Insurance (Company Provided)
-
Employee Assistance Program
-
Unlimited access to LinkedIn learning solutions
-
Matched 401(k) Retirement Savings Plan
-
Paid Time Off - the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)
-
Paid Holidays - nine (9) total per year
-
Legal Plan and Identity Theft Protection
-
Accident Insurance
-
Employee Discounts
-
Pet Insurance
-
Employee Stock Purchase Program
-
If otherwise eligible, participation in the discretionary annual bonus program
-
If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program
This Remote Position Cannot be Performed in New York City.
This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations. Most candidates are hired at a salary within the range disclosed. Salary range: $145,000 - $160,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.
In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance’s key provisions here: Concept FCO Posting 8 27 24 (lacounty.gov)
EPAM Systems, Inc. is an equal opportunity employer. We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities. We are committed to recruiting, hiring, developing and promoting employees without discrimination. As a global employer, this commitment includes complying with all laws in the countries in which we operate. Nevertheless, we believe equal employment practices should not be limited to what the law requires. Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.
At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.
#J-18808-Ljbffr$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SuggestedWork experience placementWork at officeRemote workFlexible hours- ...Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring tools (Splunk, Sumo Logic, New Relic,...Suggested
$151.2k - $204.6k
Would you like to be an engineer who builds the systems that power... ...queries every day, where latency, reliability, and quality translate... ...Development Engineer, operating as a Site Reliability Engineer, to... ...for business-critical systems, lead the engineering response to...SuggestedFlexible hours$143k - $194k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...SuggestedFull timeTemporary workWork experience placementImmediate start- ...Security team. The successful candidate will apply an engineering-based approach to solving complex security and reliability challenges, leveraging machine data analytics... ...with team members and senior technical leads to deliver quality solutionsParticipate in on-call...SuggestedFull timeInternshipSummer internshipWork at officeLocal areaRemote workFlexible hours
$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed... ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,...Work experience placementWork at officeRemote work$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwideFlexible hours- ...day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance... ..., automating away toil, and leading the technical response when production... ...years of hands-on Cloud Operations and Site Reliability Engineering, operating production...Full time
- ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry... ...: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent...Full timeWork at office
$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...adventure where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
$55k - $151.47k
...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in... ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational...Full timeH1b$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident... ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale...Remote workFlexible hours
- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers... ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them...WorldwideHome officeFlexible hours
$204k - $306k
...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,... ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and... ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week- ...certification), ISO 27001:2005 Information Security Management System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months Interview – in-person if local or Phone + Skype Minimum...Local areaWorldwide
- ...together. We are responsible for the reliability of all the company's major... ...products, services, and query engines. We serve business needs... ...effectively.- Incident Management: Lead efforts to troubleshoot and... ...emerging technologies related to site reliability and infrastructure...
$95k - $134k
...next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five... ...Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work...Temporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours- ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health...
- ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have... ...production (design → launch → on-call → reliability improvements) Led incident response and... ...and use them to prioritize work. Lead incident response for high-severity issues...
$94k - $142.3k
...to level-up your career at the company leading workforce transformation in the... ...and partner enablement, applications engineering, infrastructure, collaboration, enterprise... ...globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of...Full timeShift work- ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
- ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
- The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power... ...Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform...
$232k - $319k
...scale the service with great people and reliable, cost-effective, and efficient infrastructure... ...& tooling. What you’ll be doing Lead the Infra platform and shared services org... ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful...Permanent employmentLocal areaWorldwideFlexible hours- ...team is responsible for the reliability, scalability, and efficiency... ...building features, but about engineering the resilience and performance... ...maintain system stability.As a Site Reliability Engineer, you... ...Capacity and cost optimization: Lead initiatives in capacity planning...
- Overview Site Reliability Engineer, Compute - USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security (USDS) is a subsidiary of TikTok in the U.S. This security-first division was created to bring heightened focus and governance to data...Work experience placement
$117.2k - $313.7k
...up your career at the company leading workforce transformation in... ...Distributed Systems Software Engineer - Public Cloud (Senior/Lead/Principal... ...on our platform to be highly reliable, lightning fast, supremely... ...experience balancing live-site management, feature delivery,...Full time$120k - $170k
Sr. Manager/Manager Site Reliability Engineering Join to apply for the Sr. Manager/Manager Site Reliability Engineering role at Aritzia Sr. Manager... ...opportunity to be part of the Quality and Service Delivery team and lead the SRE team responsible for continuously improving digital...Full timeWork at officeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Seattle, WA
- lead operating engineer Seattle, WA
- lead network engineer Seattle, WA
- lead infrastructure engineer Seattle, WA
- site reliability engineer sre Seattle, WA
- site reliability engineer Seattle, WA
- official site Seattle, WA
- site services specialist Seattle, WA
- construction site safety Seattle, WA
- IT site lead Seattle, WA



