Reliability Engineer
Systems Engineering Solutions Inc Defunct
This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer . The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission‑critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities.
Location: This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA).
Requirements
System Reliability & Availability
- Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments
- Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
- Identify reliability risks and implement mitigation strategies across the system lifecycle
- Conduct capacity planning and performance modeling to ensure systems scale to meet demand
Monitoring, Observability & Alerting
- Implement and manage monitoring, logging, and tracing solutions to provide full system observability
- Define actionable alerting thresholds that minimize noise and enable rapid incident detection
- Analyze trends and metrics to proactively identify potential reliability issues
Incident Response & Problem Management
- Participate in on‑call rotations and lead incident response activities for production systems
- Coordinate troubleshooting efforts across development, infrastructure, and security teams
- Conduct post‑incident reviews (PIRs) and develop corrective and preventive action plans
- Track recurring issues and ensure root causes are resolved
Automation & Engineering Excellence
- Automate operational tasks to reduce manual intervention and operational risk
- Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR)
- Promote “automation over toil” and standardize operational workflows
Reliability‑Focused Engineering
- Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability
- Validate disaster recovery (DR) and business continuity plans; test failover mechanisms
- Support chaos engineering, fault injection testing, and resilience validation where appropriate
Collaboration & Governance
- Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives
- Document system reliability standards, runbooks, and operational procedures
- Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls)
Required Skills:
· Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree.
· Active Secret clearance at a minimum required to start
· US citizenship required
· Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services
· Experience with containerized environments (Docker, Kubernetes)
· Familiarity with CI/CD pipelines and deployment automation
· SLOs and error budgets
· Capacity modeling and performance testing
· Strong understanding of:
· Distributed systems and high‑availability architectures
· Linux/Windows system administration
· Networking fundamentals (DNS, TCP/IP, load balancing)
· Hands-on experience with:
· Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)
· Infrastructure as Code (Terraform, ARM, CloudFormation)
· Scripting or programming languages (Python, Bash, Go, PowerShell, or similar)
· Experience supporting incident management and on‑call operations
Preferred Skills
- Experience with USAF Cloud One or Platform 1.
- Experience with Zero Trust Architecture
- Cloud certifications in AWS, Azure, Google, or Oracle clouds
Benefits
SES provides a competitive salary and the following benefits:
- Medical
- Dental
- Vision
- AD&D
- STD
- LTD
- Company paid Life Insurance
- 401k with employer contribution
- Paid Time Off
- Pet Insurance
$125k - $140k
...Electrical Reliability Engineer Highland Electric Fleets' mission is to make electric fleets accessible and affordable for all, enabling communities to realize the benefits of cleaner, quieter and healthier fleets. Highland is North America's leading provider of Electrification...SuggestedTemporary workLocal areaRemote work$146k - $194k
...Reliability Engineer Costa Mesa, California, United States; Quincy, Massachusetts, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise...SuggestedFull timeWork experience placementImmediate start$150k - $195k
...Senior Reliability Engineer At WHOOP, we're on a mission to unlock human performance and healthspan. WHOOP empowers members to perform at a higher level through a deeper understanding of their bodies and daily lives. WHOOP is seeking a Senior Reliability Engineer...SuggestedFull timeWork at officeRelocation$152.4k - $254k
...Job Description Summary As a Senior Reliability Engineer, you are at the vanguard of the energy transition. You will architect the future of Power, Wind, and Electrification by driving cutting-edge systems integration, modeling, and simulation. From shaping data center...SuggestedContract workRelocation package- ...your contributions make a meaningful impact. For more information, visit POSITION OVERVIEW We are seeking a Reliability Engineer to support reliability, maintenance, and asset management activities within a regulated pharmaceutical or biotechnology...SuggestedWork at officeRemote workWork visaFlexible hours
$52.7 - $62 per hour
...Reliability Engineer This individual will provide support for Facility Operations/Engineering department sites as part of the Reliability Engineering team. The ideal candidate will be responsible for ensuring the reliability and performance of our equipment/systems...Minimum wageFlexible hours- ...multi-day technology designed to keep the electric grid secure and reliable, even during extended periods of stress. By strengthening the... .... Role Description Form Energy is seeking a Staff Reliability Engineer to improve the reliability and lifetime of our iron-air battery...Full timeRelocation package
$112k - $140k
...at a time. To those who see AI as a driver of progress, come build the future together. The Crown Is Yours As a Database Reliability Engineer, you'll support the reliability, scalability, and operational excellence of the database infrastructure powering one of the...Full timeImmediate start- Walden Robotics in Cambridge, MA seeks a Reliability & Test Engineer to define and drive reliability from design to field deployment. You will collaborate across mechanical, electrical, firmware, controls, manufacturing, and operations to identify risks and implement validation...
$130k - $150k
...and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud environments...Work at officeWork from home3 days per week$115.5k - $164.8k
...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations...Work experience placementWork at officeRemote work$90k
...SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocations: Boston/Waltham, MA | HybridGet To Know us:SS... ...out talented candidates for the position ofSite Reliability Engineer.This role is based out of one of our Boston-area offices (Boston...Ongoing contractFull timeCasual workWork at officeWorldwideFlexible hours$115k - $130k
...Sirona and its products. We are looking for a talented Site Reliaiblity Engineer II to join our team. You will manage system availability, automate operational tasks, and enhance service reliability. You will be working on a global team that will ensure system...Work experience placementWork at officeLocal areaImmediate startWorldwide$75.7k - $136.3k
...complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services...Work experience placementWork at office$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed...
- ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the... ...ideal candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based...Full time
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours- ...best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various servicesand applications that come together to deliver...Full timeFlexible hours
$160k - $200k
...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams.Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident...Temporary workWork at officeLocal areaFlexible hours3 days per week$140k - $210.9k
...Senior Site Reliability Engineer Federal Reserve Financial Services (FRFS) delivers a suite of payments services to financial institutions via FedLine® Solutions, FedNowSM, Fedwire®, National Settlement Service (NSS), FedCash®, FedACH® (Automated Clearing House), and...Full timeTemporary workPart timeWork at officeShift work$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...Temporary workImmediate startFlexible hoursShift work$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...Work experience placementWork at officeLocal areaRelocation3 days per week$121.4k - $218.6k
...complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services...Work experience placementWork at office$168k - $200k
...passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...Remote work$160k - $200k
...promQL Key Responsibilities: Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coach Perform incident...Temporary workWork at officeLocal areaFlexible hours3 days per week$134.25k - $214.8k
...Sr. Site Reliability Engineer I Boston, Massachusetts, United States Join Axon and be a Force for Good. At Axon, we're on a mission to Protect Life. We're explorers, pursuing society's most critical safety and justice issues with our ecosystem of devices and...Work experience placementWork at officeRemote workFlexible hours$160k - $200k
Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware...Local areaRemote work$169.3k - $304.7k
...specialize in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible for all... ...of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,...Work experience placementWork at officeRemote work$160k - $225k
...of thousands of users across hundreds of organizations globally. About the Role Manifold is looking for a Staff Site Reliability Engineer (SRE) to work at the intersection of AI, data infrastructure, and life sciences. In this high-impact role, you will help...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Reliability Engineer. Be the first to apply!
- senior reliability engineer Boston, MA
- sr reliability engineer Boston, MA
- reliability engineer Boston, MA
- reliability maintenance engineering technician Boston, MA
- reliability engineering manager
- senior reliability engineer
- network reliability engineer
- sr reliability engineer
- fixed equipment reliability engineer
- principal reliability engineer


