Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$210k - $230k

GovCIO

GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation, reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and is a hybrid remote/onsite position.

Responsibilities
Key Responsibilities:
Infrastructure & Automation
  • Design, deploy, and manage cloud infrastructure using Infrastructure as Code (IaC) principles
  • Develop and maintain Terraform modules for AWS and Azure environments
  • Create and manage Ansible playbooks for configuration management and application deployment
  • Implement CI/CD pipelines using GitHub Actions to automate build, test, and deployment processes
  • Implement GitOps workflows for declarative infrastructure and application delivery
  • Build self-service tools and platforms to enable development teams
Reliability & Performance
  • Establish and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Implement comprehensive monitoring, logging, and alerting solutions
  • Conduct capacity planning and performance tuning
  • Perform root cause analysis and implement preventive measures
  • Design and execute chaos engineering experiments to validate system resilience
Disaster Recovery & Business Continuity
  • Design and implement disaster recovery strategies across multi-cloud environments
  • Develop and maintain backup and restore procedures
  • Create and test business continuity plans
  • Implement automated failover mechanisms
  • Document recovery time objectives (RTO) and recovery point objectives (RPO)
Cloud Operations
  • Manage and optimize AWS services (EC2, S3, RDS, Lambda, ECS, EKS, CloudWatch, etc.)
  • Manage and optimize Azure services (VMs, Storage, SQL Database, AKS, Monitor, etc.)
  • Implement cost optimization strategies and resource tagging
  • Ensure security best practices and compliance requirements
  • Manage identity and access management (IAM) policies
Collaboration & Leadership
  • Participate in on-call rotation and incident response
  • Collaborate with development teams on architecture and design decisions
  • Mentor team members on SRE practices and tools
  • Document systems, processes, and runbooks
  • Drive continuous improvement initiatives
Qualifications
Required Education and Experience
  • Bachelor's Degree with 12+ yrs experience
Technical Skills
  • Cloud Platforms: 3+ years of hands-on experience with AWS and Azure
  • Infrastructure as Code: Expert-level proficiency with Terraform
  • Configuration Management: Strong experience with Ansible
  • Scripting: Proficiency in Python, Bash, or PowerShell
  • Containerization: Experience with Docker and Kubernetes
  • Version Control: Strong Git and GitHub workflow knowledge
  • GitOps: Experience implementing GitOps practices and workflows
  • Monitoring Tools: Experience with Prometheus, Grafana, ELK Stack, or similar
  • CI/CD: Hands-on experience with GitHub Actions, Jenkins, GitLab CI, or Azure DevOps
Core Competencies
  • Deep understanding of Microsoft/Linux systems administration
  • Strong networking knowledge (TCP/IP, DNS, load balancing, VPN)
  • Experience with database administration (PostgreSQL, MySQL, SQL Server)
  • Knowledge of security best practices and compliance frameworks
  • Understanding of microservices architecture and distributed systems
  • Experience with disaster recovery planning and execution
Soft Skills
  • Excellent problem-solving and analytical abilities
  • Strong communication skills, both written and verbal
  • Ability to work independently and in team environments
  • Customer-focused mindset with emphasis on reliability
  • Adaptability to rapidly changing technologies and requirements

Clearance Required: Active Secret with the ability to obtain and hold DEA suitability

Preferred Qualifications
  • AWS Certified Solutions Architect or SysOps Administrator
  • Azure Administrator or Solutions Architect certification
  • Certified Kubernetes Administrator (CKA)
  • HashiCorp Certified: Terraform Associate
  • GitHub Certified or demonstrated expertise with GitHub Enterprise
  • Experience with service mesh technologies (Istio, Linkerd)
  • Knowledge of observability platforms (Datadog, New Relic, Dynatrace)
  • Experience with GitOps tools and practices (ArgoCD, Flux, GitHub Actions for GitOps)
  • Familiarity with compliance frameworks (SOC 2, HIPAA, FedRAMP)
  • Previous experience in a DevOps or Platform Engineering role
Posted Salary Range

USD $210,000.00 - USD $230,000.00 /Yr.

#J-18808-Ljbffr
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Arlington, VA vacancy
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing to...  ...ensuring scalability, performance, and reliability across environments. What You’ll Do Design... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $168k - $200k

     ...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,... 
    Senior
    Remote work

    Datavant

    Washington DC
    3 days ago
  •  ..., Oracle, Cassandra, SQL Server, My SQL and Mongo DB Seniority level ~ Seniority level Mid-Senior level Employment...  ...job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago... 
    Senior
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    2 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Senior
    Remote work

    Noctua Technology

    Washington DC
    9 hours ago
  • $106.3k - $221.1k

     ...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key... 
    Senior
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    4 days ago
  • $121.4k - $218.6k

     ...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner...  ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Washington DC
    2 days ago
  • $150k - $180k

     ...redefining what’s possible in remote sensing, you belong here at Umbra. About the Job We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's... 
    Senior
    Permanent employment
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    Umbra

    Arlington, VA
    4 days ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Senior
    Work experience placement

    Samprasoft

    Washington DC
    4 days ago
  • $166k - $220k

     ...Senior Site Reliability Engineer Costa Mesa, California, United States; Washington, District of Columbia, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology.... 
    Senior
    Full time
    Work experience placement
    Immediate start

    anduril

    Washington DC
    2 days ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    5 days ago
  • $160k - $210k

     ...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale... 
    Senior
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Washington DC
    24 days ago
  • $128.5k - $190k

     ...experiences together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the...  ...applications that power a highly reliable global SaaS platform. As a Senior Site Reliability Engineer, you will play a key role in... 
    Senior
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    2 days ago
  • $207k - $284.9k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    19 days ago
  • $118k - $177k

     ...Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely . Everforth ECS is seeking talented professionals to join our successful and growing team in building the next-generation Continuous Diagnostics and Mitigation (CDM) Cyber data solution... 
    Senior
    Remote work

    ECS

    Fairfax, VA
    2 days ago
  •  ...Job Description Job Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient... 
    Senior
    Local area

    Tiger Analytics Inc.

    Washington DC
    a month ago
  • $81.1k - $187k

     .... You'll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers...  .... Responsibilities Escalation points for junior Site Reliability Engineers during complex or high-impact incidents. Manage... 
    Senior
    Temporary work
    Work experience placement
    Monday to Friday
    Flexible hours
    Shift work
    Night shift

    Hackajob

    Reston, VA
    2 days ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Senior
    Full time
    Flexible hours

    Proofpoint

    Alexandria, VA
    4 days ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,... 
    Senior
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    a month ago
  • $165k - $270k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX's Starlink technology and launch capability to support national security efforts... 
    Senior
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    1 day ago
  • $165k - $265k

     ...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we're leveraging our experience...  ...availability Mentor and train junior engineers As a senior engineer you must lead the team to technical excellence - your... 
    Senior
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    9 hours ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Senior
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    3 days ago
  • $169.3k - $304.7k

     ...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible...  ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    4 days ago
  • $84.9k - $209.5k

     ...Help ensure healthcare professionals can reliably access the applications they depend on...  ...Oracle Health is seeking a Principal Site Reliability Engineer to strengthen the reliability,...  ...valuable insights and information with senior team members, management, and beyond to... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    3 days ago
  • $112k - $179k

     ...system, network, software, and security solutions. About The Role Peraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton

    Washington DC
    4 days ago
  • $115.5k - $164.8k

     ...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    3 days ago
  • $100k - $160k

     ...Site Reliability Engineer (SRE) Team CATHEXIS elevates the government contracting experience through rapid response, deep skill, and thoughtful problem-solving and communication. Our core capabilities are our top-tier program and project management, data analytics,... 
    Temporary work

    TruVoice from Corporate Visions (Formerly Primary Intelligen...

    McLean, VA
    3 days ago
  •  ...Site Reliability Engineer Mc Lean, VA Long Term Client's Enterprise Data Machine Learning (EDML) employs innovative minds like yourself to design and develop software-systems that can meet the demand of our ever-growing customer base. Like a... 
    Immediate start

    Maintec Technologies

    McLean, VA
    1 day ago
  •  ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running...  ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Inc.

    Arlington, VA
    1 day ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    2 days ago
  •  ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise Cloud transformation... 
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!