Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$145k - $160k

EPAM Systems, Inc.

We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.

Req.#1080523852

Responsibilities

  • Disaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architectures

  • Platform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScript

  • Telemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and Firehose

  • Observability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)

  • AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routing

Requirements

  • AWS CDK & TypeScript for infrastructure provisioning and automation

  • Agent & Cluster Agent deployment/management

  • Monitors, Dashboards, Synthetics, and SLOs managed via Terraform

  • Advanced constructs: Composite monitors, cardinality management, and retention controls

  • OP2 / Vector (Observability Pipelines Worker)

  • Log routing, enrichment, and redaction

  • Splunk & GCP Pub/Sub sinks; OpenTelemetry standards

  • CloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)

  • AWS Lambda runtime monitoring & performance tuning

  • PagerDuty integration and alert-to-page wiring from Datadog

We offer

  • Medical, Dental and Vision Insurance (Subsidized)

  • Health Savings Account

  • Flexible Spending Accounts (Healthcare, Dependent Care, Commuter)

  • Short-Term and Long-Term Disability (Company Provided)

  • Life and AD&D Insurance (Company Provided)

  • Employee Assistance Program

  • Unlimited access to LinkedIn learning solutions

  • Matched 401(k) Retirement Savings Plan

  • Paid Time Off - the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)

  • Paid Holidays - nine (9) total per year

  • Legal Plan and Identity Theft Protection

  • Accident Insurance

  • Employee Discounts

  • Pet Insurance

  • Employee Stock Purchase Program

  • If otherwise eligible, participation in the discretionary annual bonus program

  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program

This Remote Position Cannot be Performed in New York City.

This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations. Most candidates are hired at a salary within the range disclosed. Salary range: $145,000 - $160,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.

In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance's key provisions here: Concept FCO Posting 8 27 24 (lacounty.gov)

EPAM Systems, Inc. is an equal opportunity employer. We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities. We are committed to recruiting, hiring, developing and promoting employees without discrimination. As a global employer, this commitment includes complying with all laws in the countries in which we operate. Nevertheless, we believe equal employment practices should not be limited to what the law requires. Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.

At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Los Angeles, CA vacancy
  • $125k - $150k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (RAPTOR)SpaceX is looking for a Site Reliability Engineer with a strong drive to solve challenging problems in the Raptor... 
    Suggested
    Permanent employment
    Temporary work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $155k - $195k

     ...you to join us on our mission of providing humankind access to the galaxy beyond our planet. About the RoleWe are seeking a Site Reliability Engineer to join our Ground Software team. As a Site Reliability Engineer, you will design, build, and operate the ground and site... 
    Suggested
    Full time
    Work at office

    Apex Technology

    Los Angeles, CA
    13 hours ago
  • $145k - $195k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (TOP SECRET CLEARANCE)As a member of the Classified IT Systems Engineering team, the Site Reliability Engineer is involved... 
    Suggested
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  •  ...enjoy unforgettable live entertainment, and we continue to lead the evolution of our industry today.We’re passionate about...  ...and brightest in technology and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing, implementing, and maintaining... 
    Full time
    Local area
    Worldwide
    Flexible hours

    AXS Group

    Los Angeles, CA
    3 days ago
  • $125k - $145k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCEAs a Site Reliability Engineer, you will design, develop, and test key aspects of an in-house... 
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $165k - $230k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    21 hours ago
  • $125k - $145k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of... 
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    21 hours ago
  • $164k - $270k

     ...individual parts.Hadrian is backed by leading investors including T. Rowe Price, Lux...  ...beyond.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs through...  ...with controls, robotics, and platform engineering teams to bake reliability in early. Review... 
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    2 days ago
  • $30.53 - $56.48 per hour

    Job Title:Associate Site Reliability EngineerRequisition ID:R027696Job Description:Job Title: Associate...  ...Associate Site Reliability Engineer helps keep Marketing Technology services...  ...Hawk’s Pro Skater, and Guitar Hero. As a leading worldwide developer, publisher and distributor... 
    Hourly pay
    Full time
    Temporary work
    Part time
    Internship
    Local area
    Worldwide
    Relocation package

    Activision

    Santa Monica, CA
    21 hours ago
  • $164k - $270k

     ...additive manufacturing, and more, while scaling our Factory-as-a-Service platform to transform how critical products are built.Backed by leading investors including JPMorgan Chase, Valor Equity Partners, Andreessen Horowitz, Founders Fund, 137 Ventures, Lux Capital, T. Rowe... 
    Permanent employment
    Full time
    Local area
    Remote work
    Flexible hours

    Hadrian

    Los Angeles, CA
    3 days ago
  • $113.3k - $205.52k

     ...and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the...  ...engineering teams to shape how their services are measured, lead the work to improve them, and use what you learn from production... 
    Work at office
    Remote work
    Worldwide
    Flexible hours

    GrabJobs

    Inglewood, CA
    3 days ago
  • $180.5k - $236.91k

     ...Hi, we're Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team...  ...team's business and technical domains such as DevOps, site reliability, and cloud best practices Lead the planning, execution and release of complex... 
    Full time
    Work at office
    Flexible hours

    Namely

    Los Angeles, CA
    4 days ago
  •  ...Role: Site Reliability Engineering (SRE) Location: Los Angeles, CA Remote position Fulltime position JD Site...  ..., Pagerduty, Powershell etc.), troubleshoot issues, and lead incident resolution. Automation & Infrastructure as Code... 
    Full time
    Remote work

    SARIAN Co

    Los Angeles, CA
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Local area

    E-Solutions

    Los Angeles, CA
    3 days ago
  • $150k - $180k

     ...Senior Cloud Reliability EngineerIrvine, California, United States; Los Angeles, California, United StatesThe Senior Cloud Reliability Engineer will be responsible for writing and integrating various open source and closed sources tools. The ideal candidate will possess... 
    Work experience placement
    Local area

    Viant

    Los Angeles, CA
    4 days ago
  • $140k - $180k

     ...Senior Site Reliability Engineer Los Angeles, CA K2 is building the largest and highest-power satellites ever flown, unlocking performance...  ...every orbit. Backed by over $1 billion in total funding from leading investors including Altimeter Capital, ICONIQ, Kleiner... 
    Permanent employment
    Shift work

    K2 Space

    Los Angeles, CA
    4 days ago
  • $181k - $265k

     ...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to...  ...and Helm Understand the importance of performant and reliable systems Education - Ideally looking for a B.A. / B.S. degree... 
    Work at office
    Immediate start
    3 days per week

    Talanto

    Los Angeles, CA
    4 days ago
  • $107.8k - $162k

     ...an expectation of a minimum of three days per week working in the office and flexibility to work remotely on the remaining days. On-site expectations may evolve over time to support business needs, with clear communication provided in advance. Job Description Operates... 
    Work at office
    Local area
    Remote work
    3 days per week

    Green Dot

    Los Angeles, CA
    3 days ago
  • $210.5k - $263.1k

     ...future of anime! About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the...  ...operational agility. Capacity Planning & Performance : Lead capacity planning and performance optimization initiatives... 
    Flexible hours

    Crunchyroll

    Los Angeles, CA
    1 day ago
  •  ...practices using OpenTelemetry, logging, and standard SLOs. Own the design and follow-up process for incident response. Mentor engineering teams on secure, performant, and resilient coding practices. Promote automated testing practices and contribute to shared... 
    Full time
    Work at office
    Immediate start
    3 days per week

    Altruist

    Los Angeles, CA
    10 days ago
  •  ...Description Job Description DISQO is a leading provider of advertising intelligence,...  ...Partner with a team of high-performing engineers and developers who are focused on delivering...  ...culture, solving for security, reliability, cost-effectiveness, and observability... 
    Full time
    Contract work
    Local area
    Flexible hours
    Shift work

    DISQO

    Los Angeles, CA
    21 days ago
  •  ...STA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity...  ...company is seeking experienced Site Reliability Engineers to take ownership of building...  ...to ensure 24/7 system availability Lead incident response efforts and conduct post... 
    Permanent employment
    Full time
    Temporary work
    Immediate start

    KēSTA I.T.

    Beverly Hills, CA
    21 days ago
  • $153.84k - $246.15k

     ...celebrate differences. We believe that belonging leads to better outcomes and a stronger...  ...and learn new ones “I can succeed as a AI Engineer Lead at Capital Group.”As a AI Engineer...  ...experiences, and agentic workflows — as reliable, production-grade systems rather than demos... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Capital Group

    Los Angeles, CA
    3 days ago
  • $194.71k - $311.53k

     ...culture designed to celebrate differences. We believe that belonging leads to better outcomes and a stronger community of associates...  ...existing skills and learn new ones “I can succeed as a AI AppSec Engineer Lead at Capital Group”As a Lead AI AppSec Engineer, you will... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Capital Group

    Los Angeles, CA
    3 days ago
  • $70 per hour

     ...The Team Leader is also responsible for diagnosing, repairing and maintaining diesel engines and vehicles to ensure optimal performance and customer satisfaction. What You Will Do Lead and support an assigned team of technicians while remaining actively involved in diagnostics... 
    Flexible hours

    fletcherjonesishiringmechanicsandtechnicians

    Beverly Hills, CA
    1 day ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Los Angeles, CA
    22 days ago
  • $114.1k - $268.18k

     ...opportunities, a world-class training facility, and leading market tools, we help our people...  ...as a key liaison between cybersecurity, engineering, infrastructure, application, and risk...  ...the bottom of our KPMG US Careers site at Benefits & How We Work . Follow... 
    Full time
    H1b
    Local area

    KPMG

    Los Angeles, CA
    6 days ago
  • $164k - $270k

     ...manufacturing, and more, while scaling our Factory-as-a-Service platform to transform how critical products are built. Backed by leading investors including JPMorgan Chase, Valor Equity Partners, Andreessen Horowitz, Founders Fund, 137 Ventures, Lux Capital, T. Rowe... 
    Permanent employment
    Full time
    Local area
    Remote work
    Relocation package
    Flexible hours

    Hadrian Automation

    Los Angeles, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!