Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Reliability Engineer

Barbaricum LLC

Senior Site Reliability Engineer

Barbaricum is a rapidly growing government contractor providing leading-edge support to federal customers, with a particular focus on Defense and National Security mission sets. We leverage more than 17 years of support to stakeholders across the federal government, with established and growing capabilities across Intelligence, Analytics, Engineering, Mission Support, and Communications disciplines. Founded in 2008, our mission is to transform the way our customers approach constantly changing and complex problem sets by bringing to bear the latest in technology and the highest caliber of talent.

Headquartered in Washington, DC's historic Dupont Circle neighborhood, Barbaricum also has a corporate presence in Tampa, FL, Bedford, IN, and Dayton, OH, with team members across the United States and around the world. As a leader in our space, we partner with firms in the private sector, academic institutions, and industry associations with a goal of continually building our expertise and capabilities for the benefit of our employees and the customers we support. Through all of this, we have built a vibrant corporate culture diverse in expertise and perspectives with a focus on collaboration and innovation. Our teams are at the frontier of the Nation's most complex and rewarding challenges. Join our team.

Barbaricum is seeking an experienced Senior Site Reliability Engineer to support the reliability, availability, automation, and operational performance of IT and cloud systems under the Military Community and Family Policy (MC&FP) Outreach and Digital Enterprise Services (MODES) contract. You will help ensure MC&FP systems are reliable, scalable, resilient, and efficiently managed through proactive monitoring, automated incident response, performance optimization, and operational dashboards that support rapid decision-making.

Responsibilities:
  • Monitor and maintain system reliability, availability, and performance across on-premises, cloud, and hybrid IT environments supporting MC&FP mission requirements.
  • Implement proactive performance monitoring, automated alerting, incident response workflows, and resilience engineering practices to reduce downtime and improve operational visibility.
  • Develop, maintain, and improve scalable automated infrastructure solutions that support reliable system operations and repeatable service delivery.
  • Implement rollback strategies, recovery approaches, and chaos engineering practices to validate resilience, reduce operational risk, and improve system stability.
  • Analyze usage patterns, capacity trends, and performance indicators to support dynamic scaling, resource optimization, and system improvement decisions.
  • Develop and maintain real-time operational dashboards, reports, and metrics that enable rapid decision-making, leadership awareness, and system optimization.
  • Respond to and resolve system outages, impairments, and service disruptions while coordinating with technical teams to minimize mission impact.
  • Conduct post-incident reviews to identify root causes, document lessons learned, and implement preventative measures that reduce recurrence.
  • Collaborate with software developers, cloud engineers, cybersecurity personnel, and operations teams to improve services, reliability patterns, deployment practices, and operational standards.
  • Create and maintain system documentation, configuration standards, operational runbooks, monitoring procedures, and service reliability guidance.
  • Automate common operations tasks to reduce manual workloads, improve consistency, and increase system efficiency.
  • Implement security best practices across operational activities, infrastructure automation, monitoring, incident response, and system administration functions.
Required Skills:
  • Expert knowledge of site reliability engineering practices, system monitoring, incident management, automation, performance tuning, and operational resilience.
  • Strong understanding of Windows and Linux administration, infrastructure operations, system configuration, service management, and troubleshooting practices.
  • Experience with automation platforms and configuration management tools such as Ansible, Puppet, Chef, or similar technologies.
  • Proficiency with scripting languages such as Python, Shell, PowerShell, or similar tools used to automate operational and infrastructure tasks.
  • Knowledge of cloud services and infrastructure across AWS, Microsoft Azure, Google Cloud, or comparable cloud environments.
  • Strong understanding of network troubleshooting, configuration, connectivity analysis, system dependencies, and performance bottleneck identification.
  • Ability to design, interpret, and maintain dashboards, alerts, metrics, logs, and operational reporting that support service health and decision-making.
  • Ability to conduct root cause analysis, post-incident reviews, and corrective action planning in complex technical environments.
  • Strong problem-solving skills and the ability to work under pressure during outages, impairments, and time-sensitive operational issues.
  • Excellent written and verbal communication skills, with the ability to explain technical findings, incident impacts, and reliability recommendations to technical and non-technical stakeholders.
Required Qualifications:
  • Bachelor's degree in Computer Science, Information Technology, Systems Engineering, Cybersecurity, or a related field; Master's degree preferred.
  • Certifications related to cloud computing, system administration, site reliability engineering, DevSecOps, or automation are beneficial.
  • 10+ years of experience in site reliability engineering, systems administration, infrastructure operations, cloud operations, DevSecOps, or a similar technical role, particularly in a government, federal, defense, or secure IT setting.
  • Demonstrated experience maintaining reliable, scalable, and efficiently managed IT systems across on-premises, cloud, or hybrid environments.
  • Experience developing automated infrastructure, operational scripts, monitoring solutions, dashboards, runbooks, and configuration standards.
  • Experience supporting incident response, system outage resolution, post-incident reviews, root cause analysis, and operational improvement initiatives.
  • Experience collaborating with development, infrastructure, cloud, cybersecurity, and program teams to improve reliability, security, and service performance.
  • DoD Secret Security Clearance.

EEO Commitment

All qualified applicants will receive consideration for employment without regard to sex, race, ethnicity, age, national origin, citizenship, physical or mental disability, medical condition, genetic information, pregnancy, family structure, marital status, ancestry, domestic partner status, sexual orientation, gender identity or expression, veteran or military status, or any other basis prohibited by law.

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Senior Reliability Engineer in Washington DC vacancy
  • $100k

     ...Senior Systems Reliability Engineer Are you passionate about applying reliability and system engineering principles to analyze and assess the resilience of future strategic weapon systems? Do you have a strong technical background in reliability engineering, safety... 
    Senior
    For contractors

    Johns Hopkins Applied Physics Laboratory

    Laurel, MD
    1 day ago
  •  ...and growing capabilities across Intelligence, Analytics, Engineering, Mission Support, and Communications disciplines. Founded...  ...challenges. Join our team.Barbaricum is seeking an experienced Senior Site Reliability Engineer to support the reliability, availability,... 
    Senior
    Contract work
    For contractors

    Barbaricum

    Washington DC
    7 hours ago
  • $113k - $141.53k

     ...opportunity for economic growth due to the availability of reliable, affordable electric power. The role serves a pivotal...  ...best-practice. A day in the life of an AES Clean Energy Senior Reliability Engineer will include, but is not limited to: Analyzing reliability... 
    Senior
    Contract work
    For contractors
    For subcontractor
    Work at office

    AES Corporation

    Arlington, VA
    1 day ago
  •  ...Senior Database Reliability Engineer (DBRE) Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve... 
    Senior

    Okta, Inc.

    Washington DC
    7 hours ago
  •  ...Arthur Platform within the Instrumentation and Electrical Engineering Department. Activities Join our dynamic US Refining & Chemicals Electrical Engineering Team as A Senior Power Distribution Reliability Engineer at our Port Arthur Platform ! What does joining... 
    Senior
    Full time
    Temporary work

    Total Energies

    Washington DC
    17 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance...  ...infrastructure vision while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build, and... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $168k - $200k

     ...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,... 
    Senior
    Remote work

    Datavant

    Washington DC
    3 days ago
  •  ...Azure, Oracle, Cassandra, SQL Server, My SQL and Mongo DB Seniority level ~ Seniority level Mid-Senior level...  ...job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago... 
    Senior
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    2 days ago
  •  ...Job Description Job Description Senior Database Reliability Engineer Role Type: Full-time Location: Remote About the Role We are looking for an experienced Database Reliability Engineer to architect, optimize, and maintain highly available PostgreSQL environments... 
    Senior
    Full time
    Remote work

    YO AI Labs

    Washington DC
    25 days ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    GovCIO

    Arlington, VA
    3 days ago
  • $106.3k - $221.1k

     ...Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key Performance... 
    Senior
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    4 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Senior
    Remote work

    Noctua Technology

    Washington DC
    7 hours ago
  • $112.2k - $178.25k

     ...Senior Safety & Reliability Engineer (Direct Placement for Client)   Primary Function: As a Safety & Reliability Engineer, you will be responsible for safety processes, methodologies and best practices, assessing and defining product architecture, examining failure... 
    Senior
    Full time
    Temporary work
    Work at office
    Monday to Friday
    Flexible hours
    Day shift

    Sigma Design

    Washington DC
    more than 2 months ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Senior
    Work experience placement

    Samprasoft

    Washington DC
    4 days ago
  • $121.4k - $218.6k

     ...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner...  ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Washington DC
    2 days ago
  • $150k - $180k

     ...redefining what’s possible in remote sensing, you belong here at Umbra. About the Job We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's... 
    Senior
    Permanent employment
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    Umbra

    Arlington, VA
    4 days ago
  • $160k - $210k

     ...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale... 
    Senior
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Washington DC
    24 days ago
  • $166k - $220k

     ...Senior Site Reliability Engineer Costa Mesa, California, United States; Washington, District of Columbia, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology.... 
    Senior
    Full time
    Work experience placement
    Immediate start

    anduril

    Washington DC
    2 days ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Senior
    Full time
    Flexible hours

    Proofpoint

    Laurel, MD
    7 hours ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    5 days ago
  • $207k - $284.9k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    19 days ago
  •  ...Job Description Job Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient... 
    Senior
    Local area

    Tiger Analytics Inc.

    Washington DC
    a month ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms... 
    Senior
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    a month ago
  • $128.5k - $190k

     ...together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure...  ...that power a highly reliable global SaaS platform. As a Senior Site Reliability Engineer, you will play a key role in designing... 
    Senior
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    2 days ago
  • $165k - $270k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX's Starlink technology and launch capability to support national security efforts.... 
    Senior
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    1 day ago
  • $165k - $265k

     ...the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we're leveraging our experience in...  ...availability Mentor and train junior engineers As a senior engineer you must lead the team to technical excellence - your... 
    Senior
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    7 hours ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Senior
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    3 days ago
  • $95k

     ...Exceptional Performance ! Chenega Services & Federal Solutions, LLC , a Chenega Professional Services company, is looking for a Reliability Engineer. In this role, the Reliability Engineer will ensure the reliability, scalability, and operational health of EDAVs Azure... 
    Contract work

    Chenega Corporation

    Washington DC
    5 days ago
  •  ...Proactively suggest and implement improvements to enhance the system's reliability, resilience, and fault tolerance.Work on automating tasks to...  ...on-call support as needed.Leads and coordinates performance engineering for medium to large initiatives.Collect and document expected... 
    Remote work

    Samprasoft

    Washington DC
    2 hours ago
  • $84.7k - $144.43k

     ...Senior Mechanical Controls Engineer - Mission Critical DLR Group is an integrated design firm with a promise to elevate the human experience through design. This fuels the work we do around the world and inspires our mission to improve the lives of our clients, our... 
    Senior
    Local area

    DLR Group

    Washington DC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Reliability Engineer. Be the first to apply!