Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$145k - $160k

EPAM Systems, Inc.

We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.

Req.#1080523852

Responsibilities
  • Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures

  • Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript

  • Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose

  • Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)

  • AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing

Requirements
  • AWS CDK & TypeScript for infrastructure provisioning and automation

  • Agent & Cluster Agent deployment/management

  • Monitors, Dashboards, Synthetics, and SLOs managed via Terraform

  • Advanced constructs: Composite monitors, cardinality management, and retention controls

  • OP2 / Vector (Observability Pipelines Worker)

  • Log routing, enrichment, and redaction

  • Splunk & GCP Pub/Sub sinks; OpenTelemetry standards

  • CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)

  • AWS Lambda runtime monitoring & performance tuning

  • PagerDuty integration and alert-to-page wiring from Datadog

We offer
  • Medical, Dental and Vision Insurance (Subsidized)

  • Health Savings Account

  • Flexible Spending Accounts (Healthcare, Dependent Care, Commuter)

  • Short-Term and Long-Term Disability (Company Provided)

  • Life and AD&D Insurance (Company Provided)

  • Employee Assistance Program

  • Unlimited access to LinkedIn learning solutions

  • Matched 401(k) Retirement Savings Plan

  • Paid Time Off - the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)

  • Paid Holidays - nine (9) total per year

  • Legal Plan and Identity Theft Protection

  • Accident Insurance

  • Employee Discounts

  • Pet Insurance

  • Employee Stock Purchase Program

  • If otherwise eligible, participation in the discretionary annual bonus program

  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program

This Remote Position Cannot be Performed in New York City.

This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations. Most candidates are hired at a salary within the range disclosed. Salary range: $145,000 - $160,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.

In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance’s key provisions here: Concept FCO Posting 8 27 24 (lacounty.gov)

EPAM Systems, Inc. is an equal opportunity employer. We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities. We are committed to recruiting, hiring, developing and promoting employees without discrimination. As a global employer, this commitment includes complying with all laws in the countries in which we operate. Nevertheless, we believe equal employment practices should not be limited to what the law requires. Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.

At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.

#J-18808-Ljbffr
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Seattle, WA vacancy
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Suggested
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    4 days ago
  •  ...Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring tools (Splunk, Sumo Logic, New Relic,... 
    Suggested

    Comtech

    Seattle, WA
    1 day ago
  • $151.2k - $204.6k

    Would you like to be an engineer who builds the systems that power...  ...queries every day, where latency, reliability, and quality translate...  ...Development Engineer, operating as a Site Reliability Engineer, to...  ...for business-critical systems, lead the engineering response to... 
    Suggested
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $143k - $194k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    4 days ago
  •  ...Security team. The successful candidate will apply an engineering-based approach to solving complex security and reliability challenges, leveraging machine data analytics...  ...with team members and senior technical leads to deliver quality solutionsParticipate in on-call... 
    Suggested
    Full time
    Internship
    Summer internship
    Work at office
    Local area
    Remote work
    Flexible hours

    F5 Networks

    Seattle, WA
    4 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    7 hours ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    1 day ago
  •  ...day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance...  ..., automating away toil, and leading the technical response when production...  ...years of hands-on Cloud Operations and Site Reliability Engineering, operating production... 
    Full time

    MangoApps

    Seattle, WA
    3 days ago
  •  ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry...  ...: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent... 
    Full time
    Work at office

    TalentDome Staffing

    Seattle, WA
    3 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    4 days ago
  •  ...adventure where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    JP Morgan Chase

    Seattle, WA
    2 days ago
  • $55k - $151.47k

     ...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in...  ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational... 
    Full time
    H1b

    PwC

    Seattle, WA
    3 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    7 hours ago
  •  ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident...  ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale... 
    Remote work
    Flexible hours

    ACI Infotech

    Seattle, WA
    2 days ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Seattle, WA
    4 days ago
  • $204k - $306k

     ...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,...  ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and...  ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Bellevue, WA
    4 days ago
  •  ...certification), ISO 27001:2005 Information Security Management System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months Interview – in-person if local or Phone + Skype Minimum... 
    Local area
    Worldwide

    Comtech LLC

    Seattle, WA
    3 days ago
  •  ...together. We are responsible for the reliability of all the company's major...  ...products, services, and query engines. We serve business needs...  ...effectively.- Incident Management: Lead efforts to troubleshoot and...  ...emerging technologies related to site reliability and infrastructure... 

    TikTok

    Seattle, WA
    1 day ago
  • $95k - $134k

     ...next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five...  ...Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work... 
    Temporary work
    For contractors
    Work experience placement
    Work at office
    Local area
    Immediate start
    Flexible hours

    DAT Freight Solutions

    Seattle, WA
    3 days ago
  •  ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health... 

    Mybridge

    Seattle, WA
    3 days ago
  •  ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have...  ...production (design → launch → on-call → reliability improvements) Led incident response and...  ...and use them to prioritize work. Lead incident response for high-severity issues... 

    Practice by Numbers

    Bellevue, WA
    3 days ago
  • $94k - $142.3k

     ...to level-up your career at the company leading workforce transformation in the...  ...and partner enablement, applications engineering, infrastructure, collaboration, enterprise...  ...globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of... 
    Full time
    Shift work

    Salesforce

    Seattle, WA
    5 days ago
  •  ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    Fairygodboss

    Seattle, WA
    3 days ago
  •  ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    Next Frontier Capital

    Seattle, WA
    4 days ago
  • The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power...  ...Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform... 

    TikTok

    Seattle, WA
    1 day ago
  • $232k - $319k

     ...scale the service with great people and reliable, cost-effective, and efficient infrastructure...  ...& tooling. What you’ll be doing Lead the Infra platform and shared services org...  ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  •  ...team is responsible for the reliability, scalability, and efficiency...  ...building features, but about engineering the resilience and performance...  ...maintain system stability.As a Site Reliability Engineer, you...  ...Capacity and cost optimization: Lead initiatives in capacity planning... 

    TikTok

    Seattle, WA
    1 day ago
  • Overview Site Reliability Engineer, Compute - USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security (USDS) is a subsidiary of TikTok in the U.S. This security-first division was created to bring heightened focus and governance to data... 
    Work experience placement

    TikTok

    Seattle, WA
    4 days ago
  • $117.2k - $313.7k

     ...up your career at the company leading workforce transformation in...  ...Distributed Systems Software Engineer - Public Cloud (Senior/Lead/Principal...  ...on our platform to be highly reliable, lightning fast, supremely...  ...experience balancing live-site management, feature delivery,... 
    Full time

    Salesforce

    Bellevue, WA
    4 days ago
  • $120k - $170k

    Sr. Manager/Manager Site Reliability Engineering Join to apply for the Sr. Manager/Manager Site Reliability Engineering role at Aritzia Sr. Manager...  ...opportunity to be part of the Quality and Service Delivery team and lead the SRE team responsible for continuously improving digital... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Aritzia

    Seattle, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!