Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Reliability Engineer

Systems Engineering Solutions Inc Defunct

This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a  Reliability Engineer . The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission‑critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities.

Location: This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA).

Requirements

System Reliability & Availability

  • Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments
  • Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
  • Identify reliability risks and implement mitigation strategies across the system lifecycle
  • Conduct capacity planning and performance modeling to ensure systems scale to meet demand

Monitoring, Observability & Alerting

  • Implement and manage monitoring, logging, and tracing solutions to provide full system observability
  • Define actionable alerting thresholds that minimize noise and enable rapid incident detection
  • Analyze trends and metrics to proactively identify potential reliability issues

Incident Response & Problem Management

  • Participate in on‑call rotations and lead incident response activities for production systems
  • Coordinate troubleshooting efforts across development, infrastructure, and security teams
  • Conduct post‑incident reviews (PIRs) and develop corrective and preventive action plans
  • Track recurring issues and ensure root causes are resolved

Automation & Engineering Excellence

  • Automate operational tasks to reduce manual intervention and operational risk
  • Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR)
  • Promote “automation over toil” and standardize operational workflows

Reliability‑Focused Engineering

  • Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability
  • Validate disaster recovery (DR) and business continuity plans; test failover mechanisms
  • Support chaos engineering, fault injection testing, and resilience validation where appropriate

Collaboration & Governance

  • Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives
  • Document system reliability standards, runbooks, and operational procedures
  • Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls)

Required Skills:

· Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree.

· Active Secret clearance at a minimum required to start

· US citizenship required 

· Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services

· Experience with containerized environments (Docker, Kubernetes)

· Familiarity with CI/CD pipelines and deployment automation

· SLOs and error budgets

· Capacity modeling and performance testing

· Strong understanding of:

· Distributed systems and high‑availability architectures

· Linux/Windows system administration

· Networking fundamentals (DNS, TCP/IP, load balancing)

· Hands-on experience with:

· Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)

· Infrastructure as Code (Terraform, ARM, CloudFormation)

· Scripting or programming languages (Python, Bash, Go, PowerShell, or similar)

· Experience supporting incident management and on‑call operations

Preferred Skills

  • Experience with USAF Cloud One or Platform 1. 
  • Experience with Zero Trust Architecture 
  • Cloud certifications in AWS, Azure, Google, or Oracle clouds 

Benefits

SES provides a competitive salary and the following benefits:

  • Medical
  • Dental
  • Vision
  • AD&D
  • STD
  • LTD
  • Company paid Life Insurance
  • 401k with employer contribution
  • Paid Time Off
  • Pet Insurance
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Reliability Engineer in Boston, MA vacancy
  • $125k - $140k

     ...Electrical Reliability Engineer Highland Electric Fleets' mission is to make electric fleets accessible and affordable for all, enabling communities to realize the benefits of cleaner, quieter and healthier fleets. Highland is North America's leading provider of Electrification... 
    Suggested
    Temporary work
    Local area
    Remote work

    Highland Electric Fleets

    Boston, MA
    1 day ago
  • $146k - $194k

     ...Reliability Engineer Costa Mesa, California, United States; Quincy, Massachusetts, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    anduril

    Quincy, MA
    18 hours ago
  • $150k - $195k

     ...Senior Reliability Engineer At WHOOP, we're on a mission to unlock human performance and healthspan. WHOOP empowers members to perform at a higher level through a deeper understanding of their bodies and daily lives. WHOOP is seeking a Senior Reliability Engineer... 
    Suggested
    Full time
    Work at office
    Relocation

    Venturefizz Product Management Community

    Boston, MA
    4 days ago
  • $152.4k - $254k

     ...Job Description Summary As a Senior Reliability Engineer, you are at the vanguard of the energy transition. You will architect the future of Power, Wind, and Electrification by driving cutting-edge systems integration, modeling, and simulation. From shaping data center... 
    Suggested
    Contract work
    Relocation package

    GE Vernova

    Boston, MA
    3 days ago
  •  ...your contributions make a meaningful impact. For more information, visit POSITION OVERVIEW We are seeking a Reliability Engineer to support reliability, maintenance, and asset management activities within a regulated pharmaceutical or biotechnology... 
    Suggested
    Work at office
    Remote work
    Work visa
    Flexible hours

    Syner-G BioPharma Group

    Boston, MA
    3 days ago
  • $52.7 - $62 per hour

     ...Reliability Engineer This individual will provide support for Facility Operations/Engineering department sites as part of the Reliability Engineering team. The ideal candidate will be responsible for ensuring the reliability and performance of our equipment/systems... 
    Minimum wage
    Flexible hours

    Cushman & Wakefield

    South Boston, MA
    18 hours ago
  •  ...multi-day technology designed to keep the electric grid secure and reliable, even during extended periods of stress. By strengthening the...  .... Role Description Form Energy is seeking a Staff Reliability Engineer to improve the reliability and lifetime of our iron-air battery... 
    Full time
    Relocation package

    Form Energy, Inc.

    Somerville, MA
    4 days ago
  • $112k - $140k

     ...at a time. To those who see AI as a driver of progress, come build the future together. The Crown Is Yours As a Database Reliability Engineer, you'll support the reliability, scalability, and operational excellence of the database infrastructure powering one of the... 
    Full time
    Immediate start

    DraftKings

    Boston, MA
    3 days ago
  • Walden Robotics in Cambridge, MA seeks a Reliability & Test Engineer to define and drive reliability from design to field deployment. You will collaborate across mechanical, electrical, firmware, controls, manufacturing, and operations to identify risks and implement validation... 

    Walden Robotics

    Cambridge, MA
    3 days ago
  • $130k - $150k

     ...and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud environments... 
    Work at office
    Work from home
    3 days per week

    CRA International

    Boston, MA
    a month ago
  • $115.5k - $164.8k

     ...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    18 hours ago
  • $90k

     ...SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocations: Boston/Waltham, MA | HybridGet To Know us:SS...  ...out talented candidates for the position ofSite Reliability Engineer.This role is based out of one of our Boston-area offices (Boston... 
    Ongoing contract
    Full time
    Casual work
    Work at office
    Worldwide
    Flexible hours

    SS&C Technologies

    Boston, MA
    1 day ago
  • $115k - $130k

     ...Sirona and its products. We are looking for a talented Site Reliaiblity Engineer II to join our team. You will manage system availability, automate operational tasks, and enhance service reliability. You will be working on a global team that will ensure system... 
    Work experience placement
    Work at office
    Local area
    Immediate start
    Worldwide

    Dentsply Sirona

    Watertown, MA
    3 days ago
  • $75.7k - $136.3k

     ...complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services... 
    Work experience placement
    Work at office

    Akamai

    Cambridge, MA
    4 days ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Cambridge, MA
    4 days ago
  •  ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed... 

    Watershed Informatics

    Cambridge, MA
    3 days ago
  •  ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the...  ...ideal candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based... 
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    4 days ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Boston, MA
    1 day ago
  •  ...best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various servicesand applications that come together to deliver... 
    Full time
    Flexible hours

    Proofpoint

    Boston, MA
    18 hours ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams.Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    18 hours ago
  • $140k - $210.9k

     ...Senior Site Reliability Engineer Federal Reserve Financial Services (FRFS) delivers a suite of payments services to financial institutions via FedLine® Solutions, FedNowSM, Fedwire®, National Settlement Service (NSS), FedCash®, FedACH® (Automated Clearing House), and... 
    Full time
    Temporary work
    Part time
    Work at office
    Shift work

    Federal Reserve Bank of Boston

    Boston, MA
    18 hours ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Boston, MA
    18 hours ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating... 
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    Boston, MA
    18 hours ago
  • $121.4k - $218.6k

     ...complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services... 
    Work experience placement
    Work at office

    Akamai

    Boston, MA
    4 days ago
  • $168k - $200k

     ...passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and... 
    Remote work

    Datavant

    Boston, MA
    18 hours ago
  • $160k - $200k

     ...promQL Key Responsibilities: Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coach Perform incident... 
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interfaces

    Somerville, MA
    1 day ago
  • $134.25k - $214.8k

     ...Sr. Site Reliability Engineer I Boston, Massachusetts, United States Join Axon and be a Force for Good. At Axon, we're on a mission to Protect Life. We're explorers, pursuing society's most critical safety and justice issues with our ecosystem of devices and... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Boston, MA
    1 day ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    a month ago
  • $169.3k - $304.7k

     ...specialize in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible for all...  ...of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,... 
    Work experience placement
    Work at office
    Remote work

    Akamai

    Cambridge, MA
    4 days ago
  • $160k - $225k

     ...of thousands of users across hundreds of organizations globally. About the Role Manifold is looking for a Staff Site Reliability Engineer (SRE) to work at the intersection of AI, data infrastructure, and life sciences. In this high-impact role, you will help... 

    Manifold AI

    Cambridge, MA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Reliability Engineer. Be the first to apply!