Lead Site Reliability Engineer
$145k - $160kEPAM Systems Inc
We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.Req.#1080523852ResponsibilitiesDisaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architecturesPlatform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScriptTelemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and FirehoseObservability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routingRequirementsAWS CDK & TypeScript for infrastructure provisioning and automationAgent & Cluster Agent deployment/managementMonitors, Dashboards, Synthetics, and SLOs managed via TerraformAdvanced constructs: Composite monitors, cardinality management, and retention controlsOP2 / Vector (Observability Pipelines Worker)Log routing, enrichment, and redactionSplunk & GCP Pub/Sub sinks; OpenTelemetry standardsCloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)AWS Lambda runtime monitoring & performance tuningPagerDuty integration and alert-to-page wiring from DatadogWe offerMedical, Dental and Vision Insurance (Subsidized)Health Savings AccountFlexible Spending Accounts (Healthcare, Dependent Care, Commuter)Short-Term and Long-Term Disability (Company Provided)Life and AD&D Insurance (Company Provided)Employee Assistance ProgramUnlimited access to LinkedIn learning solutionsMatched 401(k) Retirement Savings PlanPaid Time Off - the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)Paid Holidays - nine (9) total per yearLegal Plan and Identity Theft ProtectionAccident InsuranceEmployee DiscountsPet InsuranceEmployee Stock Purchase ProgramIf otherwise eligible, participation in the discretionary annual bonus programIf otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) ProgramThis Remote Position Cannot be Performed in New York City.This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations. Most candidates are hired at a salary within the range disclosed. Salary range: $145,000 - $160,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance's key provisions here: Concept FCO Posting 8 27 24 (lacounty.gov)EPAM Systems, Inc. is an equal opportunity employer. We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities. We are committed to recruiting, hiring, developing and promoting employees without discrimination. As a global employer, this commitment includes complying with all laws in the countries in which we operate. Nevertheless, we believe equal employment practices should not be limited to what the law requires. Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.J-18808-Ljbffr
$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...SuggestedWork at officeLocal areaRemote workWorldwideFlexible hours$143k - $194k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...SuggestedFull timeTemporary workWork experience placementImmediate start- ...Windows including patching and certificate provisioning and renewals. Able to map out existing infrastructure flows and dependencies. Leading root cause analysis meetings. Responding to incidents. Understanding of logs and monitoring tools (Splunk, Sumo Logic, New Relic,...Suggested
$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SuggestedWork experience placementWork at officeRemote workFlexible hours- ...Senior Site Reliability Engineer (SRE) Location: Seattle, hybrid - 2 times a week in the office Job Type: Full-time, direct hire Industry... ...: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent...SuggestedFull timeWork at office
$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...Work experience placementWork at officeRemote workFlexible hours$95k - $134k
...next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five... ...Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work...Temporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers... ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them...WorldwideHome officeFlexible hours
- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident... ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale...Remote workFlexible hours
- ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have... ...production (design → launch → on-call → reliability improvements) Led incident response and... ...and use them to prioritize work. Lead incident response for high-severity issues...
$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$204k - $306k
...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,... ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and... ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$55k - $151.47k
...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in... ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational...Full timeH1b- ...adventure where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...together. We are responsible for the reliability of all the company's major... ...products, services, and query engines. We serve business needs... ...effectively.- Incident Management: Lead efforts to troubleshoot and... ...emerging technologies related to site reliability and infrastructure...
- Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...
- Job Title Technical/Functional Skills: Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities: Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper...
$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content, and is trusted... ...learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS...- ...Technical Support Engineer/Site Reliability Engineer At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are...
- ...certification), ISO 27001:2005 Information Security Management System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months Interview – in-person if local or Phone + Skype Minimum...Local areaWorldwide
- ...in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the Foundations... ..., and creates the space for others to do the same. Leads with courage, knowing the possibility of greatness is...
- ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health...
- ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
$94k - $142.3k
...to level-up your career at the company leading workforce transformation in the... ...and partner enablement, applications engineering, infrastructure, collaboration, enterprise... ...globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of...Full timeShift work- ...adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup...
- ...Lead Software Engineer We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology, Infrastructure Platforms team, you are an...
$232k - $319k
...scale the service with great people and reliable, cost-effective, and efficient infrastructure... ...& tooling. What you’ll be doing Lead the Infra platform and shared services org... ...serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful...Permanent employmentLocal areaWorldwideFlexible hours- The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power... ...Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform...
- ...team is responsible for the reliability, scalability, and efficiency... ...building features, but about engineering the resilience and performance... ...maintain system stability.As a Site Reliability Engineer, you... ...Capacity and cost optimization: Lead initiatives in capacity planning...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Seattle, WA
- lead operating engineer Seattle, WA
- lead network engineer Seattle, WA
- lead infrastructure engineer Seattle, WA
- site reliability engineer sre Seattle, WA
- site reliability engineer Seattle, WA
- official site Seattle, WA
- site services specialist Seattle, WA
- construction site safety Seattle, WA
- IT site lead Seattle, WA


