Senior Site Reliability Engineer
$210k - $230kGovCIO
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation, reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and is a hybrid remote/onsite position.
Responsibilities
Key Responsibilities:
Infrastructure & Automation
- Design, deploy, and manage cloud infrastructure using Infrastructure as Code (IaC) principles
- Develop and maintain Terraform modules for AWS and Azure environments
- Create and manage Ansible playbooks for configuration management and application deployment
- Implement CI/CD pipelines using GitHub Actions to automate build, test, and deployment processes
- Implement GitOps workflows for declarative infrastructure and application delivery
- Build self-service tools and platforms to enable development teams
Reliability & Performance
- Establish and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
- Implement comprehensive monitoring, logging, and alerting solutions
- Conduct capacity planning and performance tuning
- Perform root cause analysis and implement preventive measures
- Design and execute chaos engineering experiments to validate system resilience
Disaster Recovery & Business Continuity
- Design and implement disaster recovery strategies across multi-cloud environments
- Develop and maintain backup and restore procedures
- Create and test business continuity plans
- Implement automated failover mechanisms
- Document recovery time objectives (RTO) and recovery point objectives (RPO)
Cloud Operations
- Manage and optimize AWS services (EC2, S3, RDS, Lambda, ECS, EKS, CloudWatch, etc.)
- Manage and optimize Azure services (VMs, Storage, SQL Database, AKS, Monitor, etc.)
- Implement cost optimization strategies and resource tagging
- Ensure security best practices and compliance requirements
- Manage identity and access management (IAM) policies
Collaboration & Leadership
- Participate in on-call rotation and incident response
- Collaborate with development teams on architecture and design decisions
- Mentor team members on SRE practices and tools
- Document systems, processes, and runbooks
- Drive continuous improvement initiatives
Qualifications
Required Education and Experience
- Bachelor's Degree with 12+ yrs experience
Technical Skills
- Cloud Platforms: 3+ years of hands-on experience with AWS and Azure
- Infrastructure as Code: Expert-level proficiency with Terraform
- Configuration Management: Strong experience with Ansible
- Scripting: Proficiency in Python, Bash, or PowerShell
- Containerization: Experience with Docker and Kubernetes
- Version Control: Strong Git and GitHub workflow knowledge
- GitOps: Experience implementing GitOps practices and workflows
- Monitoring Tools: Experience with Prometheus, Grafana, ELK Stack, or similar
- CI/CD: Hands-on experience with GitHub Actions, Jenkins, GitLab CI, or Azure DevOps
Core Competencies
- Deep understanding of Microsoft/Linux systems administration
- Strong networking knowledge (TCP/IP, DNS, load balancing, VPN)
- Experience with database administration (PostgreSQL, MySQL, SQL Server)
- Knowledge of security best practices and compliance frameworks
- Understanding of microservices architecture and distributed systems
- Experience with disaster recovery planning and execution
Soft Skills
- Excellent problem-solving and analytical abilities
- Strong communication skills, both written and verbal
- Ability to work independently and in team environments
- Customer-focused mindset with emphasis on reliability
- Adaptability to rapidly changing technologies and requirements
Clearance Required: Active Secret with the ability to obtain and hold DEA suitability
Preferred Qualifications
- AWS Certified Solutions Architect or SysOps Administrator
- Azure Administrator or Solutions Architect certification
- Certified Kubernetes Administrator (CKA)
- HashiCorp Certified: Terraform Associate
- GitHub Certified or demonstrated expertise with GitHub Enterprise
- Experience with service mesh technologies (Istio, Linkerd)
- Knowledge of observability platforms (Datadog, New Relic, Dynatrace)
- Experience with GitOps tools and practices (ArgoCD, Flux, GitHub Actions for GitOps)
- Familiarity with compliance frameworks (SOC 2, HIPAA, FedRAMP)
- Previous experience in a DevOps or Platform Engineering role
Posted Salary Range
USD $210,000.00 - USD $230,000.00 /Yr.
#J-18808-Ljbffr$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing to... ...ensuring scalability, performance, and reliability across environments. What You’ll Do Design...SeniorFull timeRemote workRelocationRelocation package$168k - $200k
...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,...SeniorRemote work- ..., Oracle, Cassandra, SQL Server, My SQL and Mongo DB Seniority level ~ Seniority level Mid-Senior level Employment... ...job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago...SeniorContract workRemote work
$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing...SeniorRemote work$106.3k - $221.1k
...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key...SeniorLive inWork at officeLocal area$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling...SeniorWork experience placementWork at office$150k - $180k
...redefining what’s possible in remote sensing, you belong here at Umbra. About the Job We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's...SeniorPermanent employmentWork at officeLocal areaRemote workWorldwideFlexible hours- ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment...SeniorWork experience placement
$166k - $220k
...Senior Site Reliability Engineer Costa Mesa, California, United States; Washington, District of Columbia, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology....SeniorFull timeWork experience placementImmediate start$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...SeniorTemporary workImmediate startFlexible hoursShift work$160k - $210k
...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale...SeniorWork at officeImmediate startRemote workWork from home$128.5k - $190k
...experiences together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the... ...applications that power a highly reliable global SaaS platform. As a Senior Site Reliability Engineer, you will play a key role in...SeniorTemporary workWork experience placementLocal area$207k - $284.9k
...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI...SeniorPermanent employmentLocal areaWorldwideFlexible hoursDay shift$118k - $177k
...Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely . Everforth ECS is seeking talented professionals to join our successful and growing team in building the next-generation Continuous Diagnostics and Mitigation (CDM) Cyber data solution...SeniorRemote work- ...Job Description Job Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient...SeniorLocal area
$81.1k - $187k
.... You'll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers... .... Responsibilities Escalation points for junior Site Reliability Engineers during complex or high-impact incidents. Manage...SeniorTemporary workWork experience placementMonday to FridayFlexible hoursShift workNight shift$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...SeniorFull timeFlexible hours- ...Job Description Job Description Description: Onsite in Washington, DC our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,...SeniorHourly payPermanent employmentFull timeLocal areaImmediate start
$165k - $270k
...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX's Starlink technology and launch capability to support national security efforts...SeniorTemporary workImmediate startWeekend work$165k - $265k
...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we're leveraging our experience... ...availability Mentor and train junior engineers As a senior engineer you must lead the team to technical excellence - your...SeniorTemporary workImmediate startWeekend work$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine...SeniorFull timeWork experience placementImmediate start$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,...Work experience placementWork at office$84.9k - $209.5k
...Help ensure healthcare professionals can reliably access the applications they depend on... ...Oracle Health is seeking a Principal Site Reliability Engineer to strengthen the reliability,... ...valuable insights and information with senior team members, management, and beyond to...Temporary workImmediate startFlexible hoursShift work$112k - $179k
...system, network, software, and security solutions. About The Role Peraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems...Contract workWorldwideShift work$115.5k - $164.8k
...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations...Work experience placementWork at officeRemote work$100k - $160k
...Site Reliability Engineer (SRE) Team CATHEXIS elevates the government contracting experience through rapid response, deep skill, and thoughtful problem-solving and communication. Our core capabilities are our top-tier program and project management, data analytics,...Temporary work- ...Site Reliability Engineer Mc Lean, VA Long Term Client's Enterprise Data Machine Learning (EDML) employs innovative minds like yourself to design and develop software-systems that can meet the demand of our ever-growing customer base. Like a...Immediate start
- ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running... ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident...Full timeContract workRemote workFlexible hours
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise Cloud transformation...Temporary workImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- senior operations technician Arlington, VA
- senior cloud service delivery manager Arlington, VA
- senior it service manager Arlington, VA
- senior chief engineer Arlington, VA
- sr operations manager Arlington, VA
- senior account director Arlington, VA
- senior director clinical development Arlington, VA
- sr accountant Arlington, VA
- senior IT support technician Arlington, VA
- senior financial analyst remote Arlington, VA


