Lead Site Reliability Engineer
$145k - $160kEPAM Systems, Inc.
We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep visibility across our distributed AWS footprint, and codify our monitoring infrastructure using modern tools like Terraform, AWS CDK, and TypeScript.Req.#1080523852ResponsibilitiesDisaster Recovery Observability: Design, deploy, and validate monitoring and telemetry strategies supporting our transition from single-region (us-east-1) to multi-region (us-east-2 pilot-light and future Active-Active) architecturesPlatform-as-Code (PaC): Manage and automate Datadog configurations (Monitors, Dashboards, Synthetics, SLOs, and composite alerts) and AWS infrastructure using Terraform, AWS CDK, and TypeScriptTelemetry & Ingestion Pipelines: Architect and scale high-throughput telemetry ingest pipelines utilizing AWS Lambda, Amazon S3, Amazon Kinesis, and FirehoseObservability Pipelines & Routing: Implement and maintain log routing, enrichment, and redaction workflows using OP2 / Vector (Observability Pipelines Worker) alongside Splunk, GCP Pub/Sub sinks, and OpenTelemetry (OTel)AWS-Native Monitoring & Incident Response: Configure comprehensive CloudWatch metrics and alarms for edge, ALB, CloudFront, Route 53, and VPC Lattice, alongside Lambda runtime monitoring and Datadog-to-PagerDuty alert routingRequirementsAWS CDK & TypeScript for infrastructure provisioning and automationAgent & Cluster Agent deployment/managementMonitors, Dashboards, Synthetics, and SLOs managed via TerraformAdvanced constructs: Composite monitors, cardinality management, and retention controlsOP2 / Vector (Observability Pipelines Worker)Log routing, enrichment, and redactionSplunk & GCP Pub/Sub sinks; OpenTelemetry standardsCloudWatch metrics and alarms (Edge, ALB, CloudFront, Route 53, VPC Lattice)AWS Lambda runtime monitoring & performance tuningPagerDuty integration and alert-to-page wiring from DatadogWe offerMedical, Dental and Vision Insurance (Subsidized)Health Savings AccountFlexible Spending Accounts (Healthcare, Dependent Care, Commuter)Short-Term and Long-Term Disability (Company Provided)Life and AD&D Insurance (Company Provided)Employee Assistance ProgramUnlimited access to LinkedIn learning solutionsMatched 401(k) Retirement Savings PlanPaid Time Off – the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)Paid Holidays - nine (9) total per yearLegal Plan and Identity Theft ProtectionAccident InsuranceEmployee DiscountsPet InsuranceEmployee Stock Purchase ProgramIf otherwise eligible, participation in the discretionary annual bonus programIf otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) ProgramThis Remote Position Cannot be Performed in New York City.This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations. Most candidates are hired at a salary within the range disclosed. Salary range: $145,000 - $160,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance's key provisions here: Concept FCO Posting 8 27 24 (lacounty.gov)EPAM Systems, Inc. is an equal opportunity employer. We recognize the value of diversity and inclusion in creating success for our customers, business partners, shareholders, employees and communities. We are committed to recruiting, hiring, developing and promoting employees without discrimination. As a global employer, this commitment includes complying with all laws in the countries in which we operate. Nevertheless, we believe equal employment practices should not be limited to what the law requires. Equal opportunity and inclusion are essential to motivate, empower and recognize the best in everyone.At EPAM, employment actions are based on individual qualifications, without regard to race, color, religion, creed, gender, pregnancy status, sexual orientation, gender identity, gender expression, marital or familial status, national origin, ancestry, genetics, age, disability status, veteran status, citizenship status when otherwise legally able to work, or any other characteristic protected by law.J-18808-Ljbffr
$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SuggestedWork experience placementWork at officeRemote workFlexible hours$160k - $200k
...equip their workforce with composable, connected apps, leading to higher quality work, improved efficiency, and end-to... ...evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage &...SuggestedTemporary workWork at officeLocal areaFlexible hours3 days per week- ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver...SuggestedFull timeFlexible hours
$134.25k - $214.8k
...upload. Every piece of digital evidence. Every chain of custody log that holds up in court. That's us.Axon's Platform team is the engine behind what hundreds of thousands of officers rely on every day. We're one of the world's largest blob storage customers, ingesting...SuggestedWork experience placementWork at officeRemote work$130k - $150k
...industry experts, and academics. At CRA you will be exposed to leading minds who use economic, financial, and business analysis to... ...is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable...SuggestedWork at officeWork from home3 days per week$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwideFlexible hours$55k - $151.47k
...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in... ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational...Full timeH1b$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$160k - $200k
Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware...Local areaRemote work$75.7k - $136.3k
...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages... ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company...Work experience placementWork at officeRemote work$90k
...SS&C is a leading provider of mission-critical, AI-powered technology and services empowering... ..., and technology.Job DescriptionSite Reliability EngineerLocations: Boston/Waltham, MA |... ...for the position ofSite Reliability Engineer.This role is based out of one of our Boston...Ongoing contractFull timeCasual workWork at officeWorldwideFlexible hours- ...work and education) Education Desired: Bachelor of Computer Engineering Travel Percentage: 0% We are FIS. Our technology powers... ..., etc. A mindset/desire to improve application systems reliability and automate manual support tasks, to facilitate continuous improvement...Full timeWork at officeRemote workWork from homeFlexible hours
- ...), and Check Services. We are currently leading a strategic effort to transform FRFS to... ...our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team,... ...who loves building and maintaining reliable and scalable systems, CI/CD tooling, and...Full time
- ...— Kendall Square HQ (In-Office) The Role As a Senior Site Reliability Engineer at Blitzy's Kendall Square headquarters, you will be a foundational... ...Looks Like Blitzy's platform maintains industry-leading uptime — incidents are rare, and when they occur, they are...Work at office
$140k - $210.9k
...Check Services. We are currently leading a strategic effort to... ...position will be primarily on-site with residency commutable to... ...DevOps backgrounds or software engineering backgrounds (e.g., Java Python... ...interest in operating and improving reliability of distributed production...Full timeTemporary workPart timeWork at officeShift work$86.13k - $127.19k
...Site Reliability Engineer Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like... ...to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more...Permanent employmentFull timeContract workLocal area$115k - $130k
...Dentsply Sirona and its products. We are looking for a talented Site Reliaiblity Engineer II to join our team. You will manage system availability, automate operational tasks, and enhance service reliability. You will be working on a global team that will ensure system...Work experience placementWork at officeLocal areaImmediate startWorldwide$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...Akamai powers and protects life online. Leading companies worldwide choose Akamai to...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$75.7k - $136.3k
...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages... ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company...Work experience placementWork at office- ...Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading... ...agents. Role Overview We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands...
$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role... ...future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives...Temporary workImmediate startFlexible hoursShift work- ...operational efficiency, accelerate time-to-value, and deliver better customer experiences.About The RoleWe're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'...Work at officeLocal areaRemote workWork from homeWorldwideHome officeFlexible hours
$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...Work experience placementWork at officeLocal areaRelocation3 days per week$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...- ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed...
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours- ...Onshape is hiring a Principal Software Engineer (SRE) in a hybrid role based in Boston, MA. You will lead reliability initiatives, shape strategy, and act as a technical authority to ensure the platform is fast, resilient, and scalable for customers. You will drive...
$169.3k - $304.7k
...maintaining fast, efficient, scalable, and reliable routing software and infrastructure... ...global platform. As a Principal Site Reliability Engineer - Network, you will be responsible... ...Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings...Work experience placementWork at office$200k - $250k
...The Crown Is Yours As a Principal Site Reliability Engie r , you'll shape the long-term... ...cloud and on-premise platforms, helping engineering teams build, deploy, and operate highly... ...capacity planning, and cost optimization. Lead large-scale platform initiatives across...Full timeImmediate start- ## Site Reliability Engineer (FedRAMP / Security)Boston, MA · Full-time · Senior#### About The PositionCoralogix is a modern, full-stack observability... ...with R&D to improve stability & reliability of the system* Lead the product roadmap - our product is designed for engineers...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead algorithm engineer Boston, MA
- lead engineer Boston, MA
- lead operating engineer Boston, MA
- lead network engineer Boston, MA
- lead infrastructure engineer Boston, MA
- lead web developer Boston, MA
- site reliability engineer sre Boston, MA
- site reliability engineer Boston, MA
- site reliability engineer remote Boston, MA
- official site Boston, MA


