Lead Site Reliability Engineer
$145k - $160kEPAM Systems Inc
We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.
Req.#1080523852
Responsibilities
Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures
Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript
Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose
Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)
AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing
Requirements
AWS CDK & TypeScript for infrastructure provisioning and automation
Agent & Cluster Agent deployment/management
Monitors, Dashboards, Synthetics, and SLOs managed via Terraform
Advanced constructs: Composite monitors, cardinality management, and retention controls
OP2 / Vector (Observability Pipelines Worker)
Log routing, enrichment, and redaction
Splunk & GCP Pub/Sub sinks; OpenTelemetry standards
CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)
AWS Lambda runtime monitoring & performance tuning
PagerDuty integration and alert-to-page wiring from Datadog
We offer
Medical, Dental and Vision Insurance (Subsidized)
Health Savings Account
Flexible Spending Accounts (Healthcare, Dependent Care, Commuter)
Short-Term and Long-Term Disability (Company Provided)
Life and AD&D Insurance (Company Provided)
Employee Assistance Program
Unlimited access to LinkedIn learning solutions
Matched 401(k) Retirement Savings Plan
Paid Time Off – the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)
Paid Holidays - nine (9) total per year
Legal Plan and Identity Theft Protection
Accident Insurance
Employee Discounts
Pet Insurance
Employee Stock Purchase Program
If otherwise eligible, participation in the discretionary annual bonus program
If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program
This Remote Position Cannot be Performed in New York City.
This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations. Most candidates are hired at a salary within the range disclosed. Salary range: $145,000 - $160,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.
In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance’s key provisions here: Concept FCO Posting 8 27 24 (lacounty.gov)
EPAM Systems, Inc. is an equal opportunity employer. We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities. We are committed to recruiting, hiring, developing and promoting employees without discrimination. As a global employer, this commitment includes complying with all laws in the countries in which we operate. Nevertheless, we believe equal employment practices should not be limited to what the law requires. Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.
At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.
- ...Site Reliability Engineer Comtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we... ...to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding...Suggested
- ...Site Reliability Engineer Join the innovators connecting just about anything—from families to cars to now things—on T-Mobile's biggest and best network yet. The SyncUP Things platform team has an immediate need for a Site Reliability Engineer. Responsibilities:...SuggestedContract workImmediate startRemote work
$170k - $220k
...We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack.... ...shipping process.You’ll work closely with engineers, product leads, and company leadership to ensure uptime, speed, and...Suggested- Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...Suggested
$134.25k - $214.8k
...Sr. Site Reliability Engineer I Seattle, Washington, United States At Axon, we're on a mission to Protect Life. We're explorers, pursuing society's most critical safety and justice issues with our ecosystem of devices and cloud software. Like our products, we work...SuggestedWork experience placementWork at officeRemote workFlexible hours- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...WorldwideHome officeFlexible hours
- ...Job Title: Site Reliability Engineer Location: Seattle, WA FTE Only Job Description Must Have Technical... ...security frameworks including OAuth 2.0, JWT, and SAML. • Lead API lifecycle management: design, governance, deployment, versioning...
$115.5k - $164.8k
...Site Reliability Engineer II Washington, United States Join Axon and be a Force for Good. At Axon, we're on a mission to Protect Life. We're explorers, pursuing society's most critical safety and justice issues with our ecosystem of devices and cloud software....Work experience placementWork at officeRemote work- ...Site Reliability Engineer (SRE) Location: Seattle, WA (Onsite – 4 days/week) Industry: Quick Service Restaurant (QSR) Employment Type: Contract Rate: DOE Key Responsibilities Manage and enhance the enterprise vulnerability management program using tools such...Contract work
- ...in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the Foundations... ..., and creates the space for others to do the same. Leads with courage, knowing the possibility of greatness is...
- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident... ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale...Remote workFlexible hours
- Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...
$95k - $134k
...next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five... ...10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid...Temporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content, and is trusted... ...learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS...$204k - $306k
...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,... ...Francisco Office.The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and... ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week- ...Lead Software Engineer We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology, Infrastructure Platforms team, you are an...
- ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
$194k - $267k
...something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- A leading social media platform based in Seattle is seeking a Site Reliability Engineer for its U.S. Data Security division. The role involves developing automation procedures for system efficiency, collaborating with software engineering teams, and ensuring system scalability...
- Overview Site Reliability Engineer, Compute - USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security (USDS) is a subsidiary of TikTok in the U.S. This security-first division was created to bring heightened focus and governance to data...Work experience placement
$120k - $170k
Sr. Manager/Manager Site Reliability Engineering Join to apply for the Sr. Manager/Manager Site Reliability Engineering role at Aritzia Sr. Manager... ...opportunity to be part of the Quality and Service Delivery team and lead the SRE team responsible for continuously improving digital...Full timeWork at officeRemote workFlexible hours$250k - $315k
...customers, and are trusted by the world's leading companies. We build innovative solutions... ...About the Role As part of the Infrastructure Engineering team and a people leader, you will ensure the health, stability, and reliability of Rokt's critical cloud infrastructure...Full timeWork at office$160k - $180k
...Socure is seeking a Site Reliability Engineer to own AWS infrastructure and improve Kubernetes platforms. Ideal candidates will have deep AWS expertise, strong Kubernetes fundamentals, and the ability to write production-quality code in Go or Python. The position emphasizes...$28 per hour
...Foh Lead Supervisor Location: Seattle University We are hiring immediately for full time FOH LEAD SUPERVISOR positions. Address: 901 12th Avenue, Seattle, WA 98122. Note: online applications accepted only. Schedule: Full time schedules; Sunday through Thursday...Hourly payFull timeTemporary workPart timeSummer holidayLocal areaImmediate startRemote workFlexible hoursShift work$28 per hour
...Location: Seattle University We are hiring immediately for full time FOH LEAD SUPERVISOR positions. Address : 901 12th Avenue, Seattle, WA 98122. Note: online applications accepted only . Schedule : Full time schedules; Sunday through Thursday, 7:...Hourly payFull timeTemporary workPart timeSummer holidayLocal areaImmediate startRemote workFlexible hoursShift work- AWS Senior Business Intelligence Engineer III leads analytics strategy for the ASP organization, focusing on Partner insights and analytics. The role sits at the intersection of data engineering and AI, building scalable data platforms and governance. You own end-to-end...
- Datacor's GoldSim team seeks a senior scientist/engineer to advance numerical modeling and simulation capabilities within the GoldSim framework. The role blends theory, algorithm design, and practical implementation in C++ to improve the physics-based modeling engine....
- ...strategy and drive an integrated developer experience across APIs, CLIs, consoles, and agentic workflows. You will partner with engineering, architecture, developer relations, marketing, and strategic customers to translate insights into requirements, roadmaps, and go-...
- ...Location – Local Only – Will be onsite in the Seattle HQ for 1 day per week Team culture/work environment: Team of 8 system engineers Work on retail and cloud infrastructure Work all 3 store profiles Half of the team works full time at the SSC half is...Full timeLocal areaRemote work1 day per week
- Sodexo is seeking a Culinary Supervisor to provide hands-on supervision of daily food production and kitchen operations at BENAROYA HALL. You will help ensure quality, safety, and consistency while supporting Senior Cooks and Executive Chefs in executing production plans...Shift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead operating engineer Seattle, WA
- lead infrastructure engineer Seattle, WA
- lead network engineer Seattle, WA
- lead engineer Seattle, WA
- site reliability engineer Seattle, WA
- site reliability engineer sre Seattle, WA
- site safety Seattle, WA
- website coordinator Seattle, WA
- on-site clinical research associate (traveling/remote) Seattle, WA
- site services specialist Seattle, WA

