AWS Cloud Platforms Site Reliability Engineer
$140k - $180kECS Federal
Job DescriptionEverforth ECS is seeking an experienced AWS Cloud Platforms Site Reliability Engineer to work in our Fairfax, VA office in a hybrid capacity.Everforth ECS is seeking an experienced AWS Cloud Platforms Engineer specializing in reliability and resiliency to design, operate, and continuously improve an Azure Government cloud infrastructure supporting mission-critical workloads for multiple coalition Mission Partner Network enclaves in support of the DoW community. This position will have a strong focus on infrastructure reliability, automation, High Availability (HA), Disaster Recovery (DR), and Infrastructure as Code (IaC), with Terraform serving as a primary platform for provisioning and managing cloud resources.The Cloud Platforms Engineer will work across cloud infrastructure, platform engineering, security, and application teams to ensure production environments remain reliable, scalable, recoverable, secure, and operationally sustainable. The ideal candidate combines a strong cloud architecture skill set with hands-on operational experience and an automation-first approach to infrastructure management.Design, deploy, and maintain highly available, fault-tolerant cloud infrastructure using Terraform and Infrastructure as Code principles.Support and maintain High Availability (HA) and Disaster Recovery (DR) solutions, including infrastructure redundancy, automated failover, backup and restoration, geographic resiliency, and recovery procedures.Manage the complete lifecycle of cloud infrastructure, including provisioning, configuration, operating system and platform maintenance, patching, upgrades, vulnerability remediation, and decommissioning.Support automation for infrastructure deployment and routine operational activities to reduce manual administration, configuration drift, and the potential for human error.Implement and maintain monitoring, logging, alerting, and observability capabilities to identify infrastructure degradation, capacity constraints, performance issues, and potential service disruptions before they impact users.Participate in incident response, troubleshooting, and root-cause analysis for production infrastructure events and develop corrective actions to prevent recurrence.Define and maintain infrastructure reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, recovery time objectives (RTOs), and recovery point objectives (RPOs).Develop, maintain, and regularly validate disaster recovery procedures through recovery exercises, failover testing, and infrastructure restoration testing.Evaluate cloud infrastructure capacity, performance, availability, and scalability and recommend architectural or operational improvements.Partner with application development, cybersecurity, DevOps, and platform engineering teams to establish standardized deployment patterns and resilient cloud architectures.Maintain infrastructure documentation, operational procedures, architecture diagrams, runbooks, and recovery procedures required to support production environments.Provide technical leadership and guidance regarding cloud infrastructure reliability, resiliency, automation, and operational best practices.Other duties, as assigned.Note: Salary is commensurate with skillset, qualifications, experience, and educational background.Salary Range: $140,000-180,000General Description of Benefits Job RequirementsU.S. Citizen.Active DoD Secret security clearance.High School Diploma and 9+ years of relevant experience. Alternatively, a Bachelors in a related field of study and 5+ years of experience.Required Certifications:DoD 8140 IAT Level II Security+ (or higher).AWS Certified Cloud PractitionerAWS Certified Solutions Architect – AssociateAWS Certified CloudOps Engineer – AssociateAbility to work in a hybrid capacity in Fairfax, VA (up to 3 days in office).Strong experience designing, deploying, and supporting highly available, fault-tolerant cloud infrastructure.Hands-on experience with Terraform and Infrastructure as Code (IaC) principles for provisioning and managing cloud resources.Knowledge of High Availability (HA) and Disaster Recovery (DR) architecture, including redundancy, automated failover, backup and restoration, geographic resiliency, and recovery planning.Experience managing the full cloud infrastructure lifecycle, including provisioning, configuration, patching, upgrades, vulnerability remediation, maintenance, and decommissioning.Strong infrastructure automation skills with an emphasis on reducing manual administration, configuration drift, and operational error.Experience implementing and operating monitoring, logging, alerting, and observability solutions for production infrastructure.Strong troubleshooting and diagnostic skills, including incident response, root-cause analysis, and corrective action development.Understanding of Site Reliability Engineering concepts, including SLIs, SLOs, availability targets, RTOs, and RPOs.Experience developing and validating disaster recovery procedures, including failover exercises, recovery testing, and infrastructure restoration.Ability to evaluate infrastructure capacity, performance, scalability, availability, and resiliency and recommend architectural or operational improvements.Experience working collaboratively with application development, cybersecurity, DevOps, and platform engineering teams.Ability to develop and maintain technical documentation, including architecture diagrams, operational procedures, runbooks, and recovery documentation.Strong understanding of cloud infrastructure security, vulnerability management, and operational best practices.Demonstrated ability to provide technical leadership and guidance in infrastructure reliability, resiliency, automation, and cloud operations.Experience supporting production, enterprise, regulated, or mission-critical environments is highly desirable.Strong problem-solving and decision-making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end-users to executive management).Strong problem-solving and decision-making capabilities, with a proven ability to weigh the relative costs and benefits of potential actions and identify the most appropriate solution.Highly developed interpersonal and oral/written communication skills, with the ability to effectively and professionally interact with a diverse set of stakeholders (from peers to end-users to executive management).Job DetailsJob Type: Full-timeCategory: Cloud & InfrastructureSalaried: Salaried
- ...Detail Description: The AWS Site Reliability Engineer (SRE) is responsible for the operational... ...Databricks environments built by the Platform Engineering team. You prepare and take... ...~ CloudWatch, performance tuning in cloud environments, IaC tools, Databricks management...PlatformAmazon Web ServiceCloud
- ...Job Description Job Description Site Reliability Engineer II Metro DC · Hybrid · 24/7 FedRAMP... ...and act on it, on a FedRAMP-authorised cloud platform that genuinely can't afford downtime.... ...and first-response triage. Across AWS (Commercial/GovCloud) and EKS infrastructure...PlatformAmazon Web ServiceCloudHourly payContract workFor contractorsShift workNight shiftWeekend work
- ...Site Reliability Engineer Mc Lean, VA Long Term Client... ...how we operate and scale our cloud based containerized service.... ...infrastructure built on cloud services (AWS, Kubernetes, etc) Ensure... ...resolution and improve platform resiliency Basic...PlatformAmazon Web ServiceCloudImmediate start
- .../ 100% remote within the US Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental... ...and operate our client's cloud platform that supports the entire organization... ...will have strong experience in AWS and Azure, and most of those associated...PlatformAmazon Web ServiceCloudLong term contractRemote work
$118k - $177k
...Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely . Everforth ECS is seeking talented professionals to... ...and SLIs ~3+ years of relevant experience using cloud platforms (AWS GovCloud preferred) ~3+ years of hands-on...PlatformAmazon Web ServiceCloudRemote work$109.18k - $163.77k
...provides comprehensive ad platforms for publishers, advertisers... ...around the world.Job SummaryThe Site Reliability Engineering team is responsible for... ...intersection of software engineering, cloud infrastructure, networking,... ...experience with AWS (required) and/or Oracle Cloud...PlatformAmazon Web ServiceCloudFull time- ...Site Reliability Engineer Description: We are looking for a dynamic Site Reliability Engineer... ...a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as... ...Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure...PlatformAmazon Web ServiceCloud
- ...Site Reliability Engineer Location: Occasional onsite visits to Reston VA (Zip code... ...VA. Zip code: 20190 Strong AWS SRE/Platform Engineering with Python/Java, Terraform... ...(Datadog/New Relic), Linux, and cloud-native architecture with automation...PlatformAmazon Web ServiceCloudLong term contractTemporary workH1bImmediate startRelocation
$128.5k - $190k
...Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the... ...self. The Role and Team The Site Reliability Engineering organization at Medallia brings together... ...infrastructure platforms such as AWS, OCI, or GCP. ~ Demonstrated...PlatformAmazon Web ServiceCloudTemporary workWork experience placementLocal area- ...automation. Its centralized 1EXIGER.AI platform allows organizations to manage... ...for Supplier Risk Management. Site Reliability Engineer Location: U.S. (Hybrid) This... ...systems. ~ Familiarity with cloud platforms (AWS) and secure system integration. ~...PlatformAmazon Web ServiceCloudWork at officeWork from homeFlexible hours
$152.02k - $228.03k
...provides comprehensive ad platforms for publishers,... ...responsible for ensuring the reliability, scalability, and... ...closely with data engineers and other operation sub... ....Experience with cloud platforms (AWS, GCP, Azure).Familiarity... ...on our careers site for more details.EducationBachelor...PlatformAmazon Web ServiceCloudFull time$158.5k - $230k
...Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the... ...GovCloud team and looking for a Staff Site Reliability Engineer to help scale how we operate Medallia... ...and compliant environment built on AWS GovCloud and Kubernetes. This is a...PlatformAmazon Web ServiceCloudPermanent employmentTemporary workWork experience placementWork at officeLocal areaRemote work3 days per week$230k - $250k
...GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure... ...12+ years of infrastructure/cloud engineering experience (or... ...in Kubernetes and container platforms ~ Experience working with... ...certifications AWS/Azure certifications DevOps...PlatformAmazon Web ServiceCloudRemote work$210k - $230k
...currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement,... ...optimization across multi-cloud environments. This position... ...maintain Terraform modules for AWS and Azure environments•... ...Build self-service tools and platforms to enable development teamsReliability...PlatformAmazon Web ServiceCloudCurrently hiringRemote work$130k - $180k
Job DescriptionEverforth ECS is seeking a Cloud Site Reliability Engineer (SRE)to work in our Arlington, VA office/remotely. Our Philosophy We believe... ...for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define what “reliable...PlatformAmazon Web ServiceCloudFor contractorsWork at officeRemote workShift work$99k - $225k
Site Reliability Engineer, LeadThe Opportunity: As a Lead Site Reliability Engineer... ...production systems and platforms. This role leads the design... ...operational best practices across cloud and air-gapped environments,... .... Experience with AWS CloudWatch, AWS EKS, and related...PlatformAmazon Web ServiceCloudFull timeContract workPart timeWork at officeLocal areaRemote work$150k - $180k
...and Mission Solutions (the platforms).Together, our teams develop... ...seeking an experienced Senior Site Reliability Engineer to help design, build,... ...~ Extensive experience with AWS services (EC2, S3, Lambda, VPC... ...) and deep knowledge of cloud infrastructure, networking,...PlatformAmazon Web ServiceCloudPermanent employmentWork at officeLocal areaRemote workWorldwideFlexible hours$109.18k - $163.77k
...company, provides comprehensive ad platforms for publishers, advertisers, and media... ....Job SummaryWe're looking for a Site Reliability Engineer to own cloud infrastructure, system reliability,... ...multiple regions worldwide — built on AWS and Kubernetes and managed entirely...PlatformAmazon Web ServiceCloudFull timeWorldwide- ...development, cybersecurity, data engineering and analytics, operations) ~... ...of enterprise IT operations (cloud, on-premises, hybrid... ...clearance ~ HS diploma or GED ~ AWS Associate, Azure Intermediate... ...~ experience with COTS data platform technology (Palantir,...PlatformAmazon Web ServiceCloud
- ...and AI infrastructure platform so our customers can use... ...business. Founded by engineers — and customer... ...build software for the cloud; we are "cloud maximalists... ...major cloud provider (AWS, Azure, and GCP) simultaneously... ...). Security & Reliability: Partner with...PlatformAmazon Web ServiceCloudFull time
- ...Senior Software Engineer, Snowflake Natsec... ...knowledge base on cloud infrastructure, privacy... ...cloud automation platform that include cloud... ...security, reliability, availability, and... ...platforms such as AWS, Azure, or GCP ~... ...Snowflake Careers Site for salary and benefits...PlatformAmazon Web ServiceCloudFull timeContract work
- ...cyber security role Experience with SIEM platforms, EDR solutions and network traffic analysis Understanding of cloud environments (AWS, Azure, etc.) Preferred... ...information security, computer science, engineering, or other closely related IT discipline)...PlatformAmazon Web ServiceCloud
- ...MANTECH seeks a highly technical Cloud Engineer to join our team in Herndon, VA . Join a... ...strong background in Amazon Web Services (AWS), cyber development, scripting, and automation... ...rules for AWS and other cloud-based platforms using Splunk Query Language....PlatformAmazon Web ServiceCloud
- ...will possess deep familiarity with at least one major Cloud Service Provider (CSP) and have a strong understanding... ...Service Provider (CSP), such as Amazon Web Services (AWS), Microsoft Azure, or Google Cloud Platform (GCP). - Strong understanding of cloud technology and...PlatformAmazon Web ServiceCloud
- ...Site Reliability Engineer (SRE) Reston, VA Site Reliability Engineer (SRE) Position: Site Reliability... ...creating, configuring and maintaining cloud-based applications and infrastructure... ...of applications and services: ~ AWS, EC2, Fargate, CloudFormation, RDS, ElasticCache...Amazon Web ServiceCloudContract work
$80k - $130k
...Overview: BDR Solutions LLC is seeking an experienced Cloud Cybersecurity Specialist to provide security oversight for a high-security IRS AWS and Databricks environment. This role is responsible for monitoring platform security, supporting ATO and continuous monitoring...PlatformAmazon Web ServiceCloudFull timeContract work- ...our talented team.Job Title: Lead Systems Engineer Location(s): McLean, VA Job Description... ...with Google APIs and Google Cloud Platform (GCP).Collaborate with cross-functional... ...OKTA for identity management.Exposure to AWS (nice to have, not mandatory).Proficiency...PlatformAmazon Web ServiceCloud
- ...Rivermatrix is seeking a Software Engineer-Subject Matter Expert in Chantilly,... ...and contractors. Understand the cloud environments, such as AWS or Azure. Coordinate and collaborate... ...Demonstrated experience with testing platforms such as Jest or Karma. ⢠Demonstrated...PlatformAmazon Web ServiceCloudFull timeFor contractors
- ...Team.Job Title: Systems Engineer – Lead (IAM Operations... ...a strong focus on AWS Active Directory, DevOps... ...responsibilities with Site Reliability Engineering (SRE) principles... ...manage observability platforms (Splunk, NewRelic,... ...: Ensure adherence to cloud security standards,...PlatformAmazon Web ServiceCloud
- ...Consulting Group is an AWS Consulting Partner,... ...DevOps, resulting in more reliable systems, automation at... ...provide services for cloud software development,... ...Development/software engineering or the development of... ...large scale open source platforms leveraging open source...PlatformAmazon Web ServiceCloudFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AWS Cloud Platforms Site Reliability Engineer. Be the first to apply!
- senior cloud data engineer Fairfax, VA
- cloud engineer Fairfax, VA
- aws cloud architect Fairfax, VA
- aws cloud security engineer Fairfax, VA
- cloud developer Fairfax, VA
- informatica cloud developer Fairfax, VA
- senior cloud security engineer Fairfax, VA
- senior principal cloud computing engineer Fairfax, VA
- senior cloud network engineer Fairfax, VA
- aws cloud infrastructure engineer Fairfax, VA




