Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$145k - $160k

EPAM Systems Inc

We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.Req.#1080523852ResponsibilitiesDisaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architecturesPlatform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScriptTelemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and FirehoseObservability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routingRequirementsAWS CDK & TypeScript for infrastructure provisioning and automationAgent & Cluster Agent deployment/managementMonitors, Dashboards, Synthetics, and SLOs managed via TerraformAdvanced constructs: Composite monitors, cardinality management, and retention controlsOP2 / Vector (Observability Pipelines Worker)Log routing, enrichment, and redactionSplunk & GCP Pub/Sub sinks; OpenTelemetry standardsCloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)AWS Lambda runtime monitoring & performance tuningPagerDuty integration and alert-to-page wiring from DatadogWe offerMedical, Dental and Vision Insurance (Subsidized)Health Savings AccountFlexible Spending Accounts (Healthcare, Dependent Care, Commuter)Short-Term and Long-Term Disability (Company Provided)Life and AD&D Insurance (Company Provided)Employee Assistance ProgramUnlimited access to LinkedIn learning solutionsMatched 401(k) Retirement Savings PlanPaid Time Off - the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)Paid Holidays - nine (9) total per yearLegal Plan and Identity Theft ProtectionAccident InsuranceEmployee DiscountsPet InsuranceEmployee Stock Purchase ProgramIf otherwise eligible, participation in the discretionary annual bonus programIf otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) ProgramThis Remote Position Cannot be Performed in New York City.This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations. Most candidates are hired at a salary within the range disclosed. Salary range: $145,000 - $160,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance's key provisions here: Concept FCO Posting 8 27 24 (lacounty.gov)EPAM Systems, Inc. is an equal opportunity employer. We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities. We are committed to recruiting, hiring, developing and promoting employees without discrimination. As a global employer, this commitment includes complying with all laws in the countries in which we operate. Nevertheless, we believe equal employment practices should not be limited to what the law requires. Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.J-18808-Ljbffr

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Seattle, WA vacancy
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Suggested
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    3 days ago
  • $143k - $194k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    1 day ago
  •  ...Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring tools (Splunk, Sumo Logic, New Relic,... 
    Suggested

    Comtech

    Seattle, WA
    3 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Suggested
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry...  ...: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent... 
    Suggested
    Full time
    Work at office

    TalentDome Staffing

    Seattle, WA
    2 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    5 days ago
  • $95k - $134k

     ...next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five...  ...Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work... 
    Temporary work
    For contractors
    Work experience placement
    Work at office
    Local area
    Immediate start
    Flexible hours

    DAT Freight & Analytics

    Seattle, WA
    1 day ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Seattle, WA
    1 day ago
  •  ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident...  ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale... 
    Remote work
    Flexible hours

    ACI Infotech

    Seattle, WA
    4 days ago
  •  ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have...  ...production (design → launch → on-call → reliability improvements) Led incident response and...  ...and use them to prioritize work. Lead incident response for high-severity issues... 

    Practice by Numbers

    Bellevue, WA
    1 day ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  • $204k - $306k

     ...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,...  ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and...  ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Bellevue, WA
    1 day ago
  • $55k - $151.47k

     ...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in...  ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational... 
    Full time
    H1b

    PwC

    Seattle, WA
    19 hours ago
  •  ...adventure where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    JP Morgan Chase

    Seattle, WA
    4 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    1 day ago
  •  ...together. We are responsible for the reliability of all the company's major...  ...products, services, and query engines. We serve business needs...  ...effectively.- Incident Management: Lead efforts to troubleshoot and...  ...emerging technologies related to site reliability and infrastructure... 

    TikTok

    Seattle, WA
    3 days ago
  • Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...

    Syntricate Technologies

    Seattle, WA
    15 hours ago
  • Job Title Technical/Functional Skills: Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities: Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper...

    The Dignify Solutions, LLC

    Bellevue, WA
    4 days ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content, and is trusted...  ...learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS... 

    Hive

    Seattle, WA
    5 days ago
  •  ...Technical Support Engineer/Site Reliability Engineer At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are... 

    F5

    Seattle, WA
    15 hours ago
  •  ...certification), ISO 27001:2005 Information Security Management System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months Interview – in-person if local or Phone + Skype Minimum... 
    Local area
    Worldwide

    Comtech LLC

    Seattle, WA
    5 days ago
  •  ...in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the Foundations...  ..., and creates the space for others to do the same. Leads with courage, knowing the possibility of greatness is... 

    Kaav Inc.

    Seattle, WA
    4 days ago
  •  ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health... 

    Mybridge

    Seattle, WA
    1 day ago
  •  ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    Fairygodboss

    Seattle, WA
    5 days ago
  • $94k - $142.3k

     ...to level-up your career at the company leading workforce transformation in the...  ...and partner enablement, applications engineering, infrastructure, collaboration, enterprise...  ...globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of... 
    Full time
    Shift work

    Salesforce

    Seattle, WA
    2 days ago
  •  ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup... 

    Next Frontier Capital

    Seattle, WA
    2 days ago
  •  ...Lead Software Engineer We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology, Infrastructure Platforms team, you are an... 

    Hackajob

    Seattle, WA
    3 days ago
  • $232k - $319k

     ...scale the service with great people and reliable, cost-effective, and efficient infrastructure...  ...& tooling. What you’ll be doing Lead the Infra platform and shared services org...  ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    4 days ago
  • The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power...  ...Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform... 

    TikTok

    Seattle, WA
    3 days ago
  •  ...team is responsible for the reliability, scalability, and efficiency...  ...building features, but about engineering the resilience and performance...  ...maintain system stability.As a Site Reliability Engineer, you...  ...Capacity and cost optimization: Lead initiatives in capacity planning... 

    TikTok

    Seattle, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!