Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer - Disaster Recovery & Business Continuity

$130k - $150k

CRA International

About Charles River AssociatesFor over 50 years, Charles River Associates has been a premier consulting firm that offers employees a place to learn from a diverse group of consultants, industry experts, and academics. At CRA you will be exposed to leading minds who use economic, financial, and business analysis to solve complex world problems for an impressive roster of clients, including major law firms, Fortune 100 companies, and government agencies. Through a collegial environment, formal and informal training opportunities, and a broad array of professional development resources, your experience at CRA will open doors for you throughout your career.The Information Technology (ITS) department at Charles River Associates is currently a team of more than 40 professionals dedicated to enhancing, maintaining, and developing the firm's technology infrastructure and security. The team is comprised of four functions:Service Delivery & TelecomEnterprise Application SolutionsInfrastructure, Networking and Cloud SolutionsInformation SecurityInformation Technology staff are based in the Boston, Chicago, London, Munich, New York, Oakland, San Francisco, College Station and Washington, DC offices.Mainly a Microsoft house, CRA is looking to maximize the performance of our on-premise systems and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud environments. This role blends software engineering and operations practices to reduce manual toil through automation, improve service observability, and strengthen incident response. The SRE partners closely with infrastructure, security, application, and service delivery teams to define measurable reliability targets (SLIs/SLOs), implement resilient architectures, and drive continuous improvement through blameless post-incident learning.Key ResponsibilitiesHands-on System Engineering experience with core enterprise infrastructure platforms and services, including Windows Server, VMware vSphere, VMware Site Recovery Manager (SRM), SAN technologies, and the Rubrik ecosystem, with the ability to understand dependencies, recovery workflows, and failure modes across on-premises and cloud environmentsService Ownership & Reliability Targets: Partner with service owners to define and maintain service level indicators (SLIs) and service level objectives (SLOs) for availability, latency, and performance; track error budgets and reliability risk.Observability: Implement and continuously improve monitoring, logging, alerting, and dashboards to provide actionable, symptom-based signals and reduce mean time to detect/respond (MTTD/MTTR).Blameless Postmortems & Continuous Improvement: Facilitate post-incident reviews, identify root causes and contributing factors, and drive remediation items to completion; standardize learnings into runbooks and operational practices.DR Testing Program Build-Out: Design and launch a scalable DR testing program (scope, test types, cadence, success criteria, and evidence capture) in partnership with application, infrastructure, and security teams; maintain runbooks and lead regular tabletop and technical recovery exercises to validate RTO/RPO assumptions and improve recoverability.DR Readiness: Contribute to reliability architecture and disaster recovery readiness for key services, including dependency mapping, recovery testing inputs, and validation of recovery procedures.Cross-Functional Collaboration: Work day-to-day with infrastructure, network, cloud, security, and application teams to improve operational excellence, reliability culture, and shared ownership of production outcomes.Relevant Skills & ExperienceExperience operating and improving reliability of production services (on-prem and/or cloud), including incident response, operational readiness, and service ownershipWorking knowledge of SRE concepts and practices such as SLIs/SLOs, error budgets, monitoring/alerting strategy, and blameless postmortemsExperience with observability tooling and practices (logs, metrics, tracing, dashboards) and using data to drive reliability and performance improvementsExperience with disaster recovery orchestration and recovery testing using VMware Site Recovery Manager (SRM) and Azure Site Recovery (ASR) (or similar public cloud DR services)Proven experience building and operating a DR testing program, including dependency mapping, test planning, coordination across stakeholders, execution of tabletop and technical failover tests, documentation of results, and tracking remediation actions to closureStrong cross-functional communication and teamwork skills; comfortable partnering with engineering, security, and operations teams to drive shared outcomesAbility to document and standardize operational procedures (runbooks), participate in on-call rotations, and manage multiple priorities in a fast-moving environmentCareer Growth and Benefits CRA’s robust skills development programs, including a commitment to offering 100 hours of training annually through formal and informal programs, encourage you to thrive as an individual and team member. Beginning with research and analysis skill building, training continues with technical training, presentation skills, internal seminars, and career mentoring and performance coaching from an assigned senior colleague. Additional leadership and collaboration opportunities exist through internal firm development activities.We offer a comprehensive total rewards program including a superior benefits package, wellness programming to support physical, mental, emotional and financial well-being, and in-house immigration support for foreign nationals and international business travelers.Work Location FlexibilityCRA creates a work environment that enables our colleagues to benefit from being together in the office to best deliver on our promise of career growth, mentorship and inclusivity. At the same time, we recognize that individuals realize a range of benefits when working from home periodically. We currently expect that individuals spend at least 3 to 4 days a week working in the office (which may include traveling to another CRA office or to client meetings), with specific days determined in coordination with your practice or team.Our Commitment to Equal Employment OpportunityCharles River Associates is an equal opportunity employer (EOE). All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, disability, status as a protected veteran, or any other protected characteristic under applicable law.Salary and other compensationA good-faith estimate of the annual base salary range for this position is $130,000 - $150,000. Stating pay within this range may vary based on factors such as education level, experience, skills, geographic location, market conditions, and other qualifications of the successful candidate. This position may be eligible for additional bonus incentive compensation.CRA offers a comprehensive benefits package, subject to eligibility requirements, which may include: medical, dental, and vision insurance; 401(k) retirement plan with employer match; life and disability insurance; paid time off (vacation, sick leave, holidays); paid parental leave; wellness programs and employee assistance resources; and commuter benefits.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer - Disaster Recovery & Business Continuity in Boston, MA vacancy
  •  ...Job Title: Site Reliability Engineer Location: Remote with Quarterly visits to Chennai, Tamil...  ...to developing and maintaining our continuous delivery platform and container registry...  ...in SQL scripting, PowerShell, and disaster recovery strategies. Knowledge of SIEM... 
    Suggested
    Full time
    Remote work

    Saviance

    Boston, MA
    2 days ago
  • $81.1k - $187k

     ...We are looking for a Site Reliability Engineer 3 to support mission-critical...  ...contributing to capacity planning, disaster recovery, and operational...  ...to contribute to business development decisions (e....  ...sharing and best practices. Continuous Learning: Embraces continuous... 
    Suggested
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Boston, MA
    19 hours ago
  • $84.9k - $209.5k

     ...professionals can reliably access the...  ...seeking a Principal Site Reliability Engineer to strengthen the...  ...in clinical and business workflows. In...  ...validation, and recovery. Reduce...  ..., failover, and disaster recovery through...  ...strategies. Continuous Learning: ~Pursues... 
    Suggested
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Boston, MA
    19 hours ago
  • $145k - $160k

     ...specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and...  ...critical to our multi-region disaster recovery roadmap. You will architect and implement...  ...; the demand for the role; and overall business and labor market considerations. Most candidates... 
    Suggested
    Temporary work
    Remote work
    Flexible hours

    EPAM Systems Inc

    Boston, MA
    1 day ago
  • $125k - $145k

     ...lasting impact for our investors, teams, businesses, and the communities in which we...  ....The Senior Application Support Engineer ensures stability, performance, and...  ...(ServiceNow)Contribute to business continuity and disaster recovery planning and testingQualificationsRequired... 
    Suggested
    Full time

    Bain Capital

    Boston, MA
    a month ago
  •  ...growth by fostering business development, infrastructure...  ...the Director of Engineering Innovation,...  ...delivery, and enhance the reliability, scalability, and security...  ...capabilities. Drive continuous improvement...  ...continuous improvement of disaster recovery, Continuity of... 
    Full time
    Contract work
    Part time
    Work experience placement
    Work at office
    Work from home
    Monday to Friday

    State of Massachusetts, USA

    Boston, MA
    a month ago
  • $102k - $238.1k

     ...Payments Platform Engineer Position...  ...engineering, operations, business stakeholders, and...  ...the organization continues its...  ...to ensure secure, reliable, and scalable solutions...  ...resiliency, and disaster recovery design. . Proven...  ...directed to our site that is dedicated... 
    Work at office
    Local area
    Boston, MA
    13 days ago
  • $145k - $160k

     ...experienced Cloud Platform Engineer to help design,...  ...automation, and reliability across the...  ...platforms that support business and application...  ...availability, disaster recovery, security, and...  ...recovery, and business continuity solutions....  ...vision plans, plus on-site gym access,... 
    Full time
    Summer work
    Immediate start
    Remote work
    Overseas

    Grand Circle Travel

    Boston, MA
    a month ago
  •  ...systems scale securely and reliably is core to this mission.As a Platform Engineer II on the Application...  ...experience in DevOps, Site Reliability...  ...resource isolation, and disaster recovery.Experience with containerized...  ...as experience. As we continue to build a diverse and... 
    Work at office
    Relocation

    WHOOP

    Boston, MA
    15 days ago
  •  ...As a Senior Platform Engineer , you’ll be at the heart of our...  ...systems are performant, stable, reliable, and secure. You’ll own key...  ...everything you can, and continuously improve how we deliver, monitor...  ..., and improve elasticity, disaster recovery, and high-availability... 

    7AI, Inc.

    Boston, MA
    2 days ago
  •  ...residents, restored Boston Harbor, and continue to invest in protecting vital public resources...  ..., processes and standards. Analyzes business requirements and transforms them into...  ...and execution of IT Continuity and Disaster Recovery Plans. Performs related duties as... 
    Work at office
    Local area
    Remote work
    Monday to Friday

    Massachusetts Water Resources Authority

    Chelsea, MA
    8 days ago
  • $80k - $150k

     ...experienced Senior Software Engineer to join our...  ...resilient systems that meet business needs. The ideal...  ...new functionality and continuous improvements. Optimize...  ...performance, reliability, resiliency, and security...  ...infrastructure upgrades, disaster recovery testing, and cross-... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    1 day per week

    Wellington Management

    Boston, MA
    1 day ago
  •  ...Technology group delivers secure, reliable technology solutions that...  ...and supporting seamless business operations. The ideal...  ...environments as needed. Continuously monitor Kafka clusters for...  ...fixes. Develop and maintain disaster recovery plans, conducting regular testing... 
    Remote work
    Flexible hours

    Dtcc

    Boston, MA
    19 hours ago
  • $115.5k - $164.8k

     ...company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will...  ...previously required human intervention with reliable, tested automation. You will also...  ..., transferable skills, work experience, business needs, geographic market, and often a combination... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    19 hours ago
  • $115k - $130k

     ...Dentsply Sirona and its products. We are looking for a talented Site Reliaiblity Engineer II to join our team. You will manage system availability, automate operational tasks, and enhance service reliability. You will be working on a global team that will ensure system... 
    Work experience placement
    Work at office
    Local area
    Immediate start
    Worldwide

    Dentsply Sirona

    Watertown, MA
    3 days ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Cambridge, MA
    4 days ago
  • $90k

     ...technology.Job DescriptionSite Reliability EngineerLocations: Boston/...  ...position ofSite Reliability Engineer.This role is based out of...  ...Flexibility: Hybrid Work Model and Business Casual Dress Code, including...  ...personal drive to learn and continuously improve yourself, and the... 
    Ongoing contract
    Full time
    Casual work
    Work at office
    Worldwide
    Flexible hours

    SS&C Technologies

    Boston, MA
    1 day ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Cambridge, MA
    4 days ago
  • $160k - $200k

     ...observability best practices, SLIs/SLOs, and reliability culture across engineering teams.Contributing to and...  ...related knowledge & skills, experience, business needs, geographical location,...  ...test as a condition of employment or continued employment. An employer who... 
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    19 hours ago
  •  ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various servicesand applications that come together to deliver... 
    Full time
    Flexible hours

    Proofpoint

    Boston, MA
    19 hours ago
  •  ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed... 

    Watershed Informatics

    Cambridge, MA
    3 days ago
  •  ...our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team,...  ..., and the implementation and driving of continuous improvement initiatives. You will work...  ...someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and... 
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    4 days ago
  • $136.2k - $214.01k

     ...Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the...  ..., and we celebrate exceptional execution that ensures we continue to defend data and protect people. Proofpoint is an... 
    Full time
    Flexible hours

    Proofpoint

    Boston, MA
    1 day ago
  • $140k - $210.9k

     ...position will be primarily on-site with residency commutable to...  ...backgrounds or software engineering backgrounds (e.g., Java Python...  ...in operating and improving reliability of distributed production systems...  ...and driving of continuous improvement initiatives.... 
    Full time
    Temporary work
    Part time
    Work at office
    Shift work

    Federal Reserve System

    Boston, MA
    2 days ago
  • $134.25k - $214.8k

     ...Sr. Site Reliability Engineer I Boston, Massachusetts, United States Join Axon and be a Force...  ...engineers to promote self-service. Continually seek improvement within the entire platform...  ...skills, work experience, business needs, geographic market, and often a... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Boston, MA
    1 day ago
  • $121.4k - $218.6k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Cambridge, MA
    4 days ago
  • $168k - $200k

     ...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and... 
    Remote work

    Datavant

    Boston, MA
    19 hours ago
  • $160k - $200k

     ...observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and...  ...knowledge & skills, experience, business needs, geographical location,...  ...test as a condition of employment or continued employment. An employer who violates... 
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interfaces

    Somerville, MA
    1 day ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating... 
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    Boston, MA
    19 hours ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer - Disaster Recovery & Business Continuity. Be the first to apply!