Reliability Engineer
SES
This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer . The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission‑critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities.
Location
This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA).
Requirements
System Reliability & Availability
- Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments
- Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
- Identify reliability risks and implement mitigation strategies across the system lifecycle
- Conduct capacity planning and performance modeling to ensure systems scale to meet demand
Monitoring, Observability & Alerting
- Implement and manage monitoring, logging, and tracing solutions to provide full system observability
- Define actionable alerting thresholds that minimize noise and enable rapid incident detection
- Analyze trends and metrics to proactively identify potential reliability issues
Incident Response & Problem Management
- Participate in on‑call rotations and lead incident response activities for production systems
- Coordinate troubleshooting efforts across development, infrastructure, and security teams
- Conduct post‑incident reviews (PIRs) and develop corrective and preventive action plans
- Track recurring issues and ensure root causes are resolved
Automation & Engineering Excellence
- Automate operational tasks to reduce manual intervention and operational risk
- Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR)
- Promote "automation over toil" and standardize operational workflows
Reliability‑Focused Engineering
- Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability
- Validate disaster recovery (DR) and business continuity plans; test failover mechanisms
- Support chaos engineering, fault injection testing, and resilience validation where appropriate
Collaboration & Governance
- Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives
- Document system reliability standards, runbooks, and operational procedures
- Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls)
Required Skills
- Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree
- Active Secret clearance at a minimum required to start
- US citizenship required
- Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services
- Experience with containerized environments (Docker, Kubernetes)
- Familiarity with CI/CD pipelines and deployment automation
- SLOs and error budgets
- Capacity modeling and performance testing
- Strong understanding of:
- Distributed systems and high‑availability architectures
- Linux/Windows system administration
- Networking fundamentals (DNS, TCP/IP, load balancing)
- Hands‑on experience with:
- Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)
- Infrastructure as Code (Terraform, ARM, CloudFormation)
- Scripting or programming languages (Python, Bash, Go, PowerShell, or similar)
- Experience supporting incident management and on‑call operations
Preferred Skills
- Experience with USAF Cloud One or Platform 1.
- Experience with Zero Trust Architecture
- Cloud certifications in AWS, Azure, Google, or Oracle clouds
Benefits
SES provides a competitive salary and the following benefits:
- Medical
- Dental
- Vision
- AD&D
- STD
- LTD
- Company paid Life Insurance
- 401k with employer contribution
- Paid Time Off
- Pet Insurance
$125k - $140k
...nation’s largest electric school bus fleets, Highland delivers reliable, cost-effective solutions that support local communities and drive... ...of transportation. Summary: As an Electrical Reliability Engineer, you will own the field-side technical integrity of the...SuggestedTemporary workLocal areaRemote work- ...phase balancing, and grounding meet the standards required for reliable charger and vehicle operation. Assess as-built construction... ...remediation. Work with Highland's Construction Electrical Engineering and Procurement teams to build field feedback and reliability...SuggestedTemporary workRemote work
$52.7 - $62 per hour
...TitleReliability EngineerJob Description SummaryThis individual will provide support for Facility Operations/Engineering department sites as part of the Reliability Engineering team. The ideal candidate will be responsible for ensuring the reliability and performance of...SuggestedMinimum wageFull timeFlexible hours$166k - $220k
...autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.ABOUT THE TEAMThe Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems...SuggestedFull timeWork experience placementImmediate start- ...your talents are valued, and your contributions make a meaningful impact. POSITION OVERVIEW Syner-G is seeking a Facilities Engineer with 2-6 years of experience supporting building and plant operations within laboratory, manufacturing, and utility environments....SuggestedFor contractorsFor subcontractorWork at officeRemote workWork visaFlexible hoursShift work
- ...Arcadis is seeking a Reliability Engineer to support GxP Facility Operations, Engineering, and Capital Project Management teams within a regulated pharmaceutical and biotechnology environment. This individual will be responsible for implementing maintenance and reliability...Relocation
$152.4k - $254k
...Job Description Summary As a Senior Reliability Engineer, you are at the vanguard of the energy transition. You will architect the future of Power, Wind, and Electrification by driving cutting-edge systems integration, modeling, and simulation. From shaping data center...Contract workRelocation package$150k - $195k
...Senior Reliability Engineer At WHOOP, we're on a mission to unlock human performance and healthspan. WHOOP empowers members to perform at a higher level through a deeper understanding of their bodies and daily lives. WHOOP is seeking a Senior Reliability Engineer...Full timeWork at officeRelocation- ...Role Description Form Energy is seeking a Staff Reliability Engineer to improve the reliability and lifetime of our iron‑air battery cell. As part of our Reliability team, you will work closely with the Cell and Electrode Engineering teams to develop test plans, analyze...Full timeRelocation package
$166k - $220k
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry...Full timeWork experience placementImmediate start$112k - $140k
...shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together.As a Database Reliability Engineer, you'll support the reliability, scalability, and operational excellence of the database infrastructure powering one of the...Full timeImmediate start$115.5k - $164.8k
...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations...Work experience placementWork at officeRemote work$85k - $149k
...manufacture faster than ever before.We’re a team of hands-on builders, engineers, and innovators reinventing how the world makes physical... ...future of fabrication, come build it with us.Your Impact: The Reliability Engineering Team is the company's independent voice on product...Full timeWork at officeWorldwideFlexible hours$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwideFlexible hours$130k - $150k
...and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud environments...Work at officeWork from home3 days per week- ...best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver...Full timeFlexible hours
$160k - $200k
...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident...Temporary workWork at officeLocal areaFlexible hours3 days per week$90k - $110k
...SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocations: Boston/Waltham, MA | HybridGet To Know us:SS... ...out talented candidates for the position ofSite Reliability Engineer. This role is based out of one of our Boston-area offices (Boston...Ongoing contractFull timeCasual workWork at officeWorldwideFlexible hours$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and...Work experience placementWork at officeRemote workFlexible hours$136.2k - $214.01k
...class outcomes Visionary in future focused problem-solving Exceptional in execution and impact. The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to deliver...Flexible hours- ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the... ...ideal candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based...Full time
- ...work and education) Education Desired: Bachelor of Computer Engineering Travel Percentage: 0% We are FIS. Our technology powers... ..., etc. A mindset/desire to improve application systems reliability and automate manual support tasks, to facilitate continuous improvement...Full timeWork at officeRemote workWork from homeFlexible hours
$75.7k - $136.3k
...complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services...Work experience placementWork at officeRemote work$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that... ...maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering...Local areaWorldwideFlexible hours$160k - $200k
Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware...Local areaRemote work$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b- ...Anduril Industries in Quincy, MA seeks a Field Reliability Engineer for Undersea Reconnaissance & Strike. You will debug real-world field anomalies, reproduce with recorded data, and ship fixes across Rust, C++, Python, or Go. You’ll validate in simulated or hardware...
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...Work experience placementWork at officeLocal areaRelocation3 days per week$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...Temporary workImmediate startFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Reliability Engineer. Be the first to apply!
- reliability maintenance engineering technician Boston, MA
- reliability engineer Boston, MA
- database reliability engineer
- reliability maintenance engineering technician
- principal reliability engineer
- fixed equipment reliability engineer
- reliability engineer
- reliability engineering manager
- maintenance & reliability engineer
- sr reliability engineer


