Site Reliability Engineer
ClearanceJobs
Senior Site Reliability EngineerThe Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability of cloud native systems. Our SREs don't just manage infrastructure; they build it using Infrastructure as Code (IaC), monitor it through advanced observability stacks, and protect it by engineering for failure. We work closely with clients to bridge the gap between development and operations. We are seeking a highly experienced and autonomous Senior Site Reliability Engineer (SRE) to join our dynamic team. As a technical leader, you will define the strategy and apply advanced software engineering principles to operations, focusing on the architecture, reliability, and long-term performance of large-scale production systems. You will play a crucial role in reducing toil through automation, defining and monitoring Service Level Objectives (SLOs), and implementing best practices for system stability and incident response. This role requires working with modern cloud technologies to ensure the high availability and efficiency of applications and infrastructure. Location: Primarily Remote. Candidates must be based in CA or DC Metro Area for proximity to project and client teams. Security Clearance Requirement: Applicants must be US citizens and eligible to obtain and maintain an active Secret security clearance or above.Key ResponsibilitiesSite Reliability EngineeringDrive the definition and adoption of SLIs and SLOs across multiple services or entire platforms, ensuring alignment with business goals.Design and architect Infrastructure as Code (IaC) solutions for large-scale, complex environments, establishing standards and best practices.Implement and manage containerized and serverless architectures using Docker, Kubernetes, and cloud-native services, focusing on performance and error budgets.Build and maintain reliable and self-healing CI/CD pipelines to automate deployments and improve development workflows.Toil Reduction and Incident ManagementImplement and refine comprehensive monitoring, alerting, and logging to detect and address performance and availability issues proactively.Lead the strategic effort to eliminate toil, identifying and championing major automation projects that deliver significant organizational efficiency.Lead high-severity incident response and coordinate blameless postmortems for major outages, driving the resulting remediation and systemic improvements.Testing and Service ResiliencyImplement cloud security best practices, including identity and access management (IAM), encryption, and compliance controls.Proactively identify and address system weaknesses and ensure performance under stress.Support disaster recovery and high availability strategies through backup and failover planning.Collaboration and Knowledge SharingServe as a primary SRE liaison for development teams, influencing application architecture and design to meet reliability and scalability targets from inception.Create and maintain documentation for cloud architectures, deployment processes, and best practices.Contribute to internal knowledge-sharing initiatives, ensuring continuous learning within the team.Stakeholder CommunicationAct as a subject matter expert and trusted advisor to clients and internal leadership on cloud infrastructure, reliability strategy, and Service Level Agreement (SLA) negotiations.Act on client feedback to refine and enhance cloud solutions.Conduct training and knowledge-sharing sessions to help clients manage their cloud environments effectively.Continuous Learning and InnovationStay updated on the latest developments in cloud infrastructure and technology trends.Drive innovation by proposing and implementing new techniques and technologies.Qualifications5+ years of experience in site reliability engineering, cloud engineering, or related fields.Strong software engineering skills with an emphasis on writing clean, modular, and maintainable code, specifically for automation and system management.Deep experience with Infrastructure as Code (IaC) tools like Terraform or CloudFormation.Deep experience with containerization and orchestration tools like Docker and Kubernetes.Deep knowledge of networking concepts, cloud security best practices, and identity management.Experience with programming or scripting languages such as Python, Bash, or Go.Experience with CI/CD pipelines and DevOps methodologies.Strong problem-solving skills and the ability to troubleshoot complex cloud environments.Demonstrated ability to influence technical decision-making across organizational boundariesPreferred qualifications:Bachelor's or advanced degree in Computer Science or a related field.Any of the below cloud certifications: Google Cloud Professional Cloud Architect Google Cloud Professional Cloud DevOps Engineer AWS Certified Solutions Architect AWS Certified Developer AWS Certified SysOps Administrator CompTIA Security+ certification or an equivalent DoD 8140/8570 IAT Level II baseline certification.
$81.1k - $187k
...infrastructure and/or service according to terms for reliability and functionality.- Assists team members... ...deployments.- Gains basic knowledge of site reliability trends and shares relevant... ...are seeking a skilled Site Reliability Engineer to design, build, operate, and automate...SuggestedTemporary workImmediate startFlexible hoursShift work$133k - $190k
Site Reliability Engineer needed for a full time opportunity with SOC's direct client based in Herndon, VA. Direct Hire Role **Due to federal requirements, candidates must hold and possess an Active DOW TS/SCI security clearance to be considered for this role.** SOC is...SuggestedFull time$119.8k - $234.7k
...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany:... ...opportunity for a Senior Site Reliability Engineer (SRE) to join the Azure Silver and Sovereign...SuggestedOngoing contractLocal area3 days per week$135.8k - $183.8k
...dynamic and flexible work environment with competitive benefits and the ability to grow your career.We are looking for a Site Reliability Engineer to support our team responsible for building, managing, maintaining, deploying, and securing mission-critical services to...SuggestedWork at officeFlexible hours$80k - $133k
...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and...SuggestedPermanent employmentFull timeContract workRemote workFlexible hours- We're seeking a skilled and proactive Site Reliability Engineer to join our team, ensuring the stability, security, and efficiency of our technological resources as we deliver cutting-edge AI solutions to the government. This is a fully remote position for candidates in...Remote work
$91.4k - $187k
Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible...Temporary workFlexible hoursShift workWeekend work$146k - $194k
...focused on positioning Anduril as a lead provider of specialized engineering and products for Intelligence Community (IC) customers. We... ...pressing national security requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure the health,...Full timeWork experience placementImmediate startRemote work$109.18k - $163.77k
...channels. As a global company, we have offices in nine countries and can insert advertisements around the world.Job SummaryThe Site Reliability Engineering team is responsible for managing the critical infrastructure that powers FreeWheel's Streaming Hub platform. Streaming...Full time$84.24k - $142.48k
OverviewJoin us to work collaboratively with our talented team of dynamic and passionate engineers to deliver capabilities that enable our customers to make a difference. You'll deploy and operate ArcGIS Velocity and ArcGIS Workflow Manager SaaS solutions. You will also...WorldwideFlexible hours$87.1k - $157.45k
...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across... ...to come in and help us build systems that stay reliable when things get complicated. We need a Site Reliability Engineer who has experience building, deploying...Local areaImmediate startWork from homeFlexible hours- ...Senior Site Reliability EngineerThe Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long...Remote work
$158.5k - $230k
...extraordinary experiences together. Bring your whole self. The Role and Team We are growing our GovCloud team and looking for a Staff Site Reliability Engineer to help scale how we operate Medallia's US public-sector cloud platform. You will support federal agencies and other...Permanent employmentTemporary workWork experience placementWork at officeLocal areaRemote work3 days per week$128.5k - $190k
...We empower exceptional people to create extraordinary experiences together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure and applications that power a highly reliable global SaaS...Temporary workWork experience placementLocal area$121.5k - $264.1k
...sharing guidance on practices and terms for reliability and functionality.- Supervises team... ...developing and maintaining knowledge of site reliability trends and sharing valuable... ...Experience:9 years of experience in software engineering, infrastructure management, or related...Temporary workImmediate startFlexible hours$102k - $234.6k
...members in designing and architecting infrastructure and service for reliability and functionality. Provides day-to-day direction to help... ...to experiment with new technology, execute improvements, build site reliability knowledge, and provide clear data.Only Oracle brings...Temporary workImmediate startFlexible hours$84.9k - $209.5k
.... You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers... ...posted.Career Level - IC4Escalation points for junior site reliability engineers during complex or high-impact incidents.Manage and...Temporary workMonday to FridayFlexible hoursShift workNight shift$128.83k - $193.25k
...you will be responsible for ensuring the reliability, scalability, and performance of our data systems. Working closely with data engineers and other operation sub-teams, you will manage... ...and benefits summary on our careers site for more details.EducationBachelor's DegreeWhile...Full time- ...Required U.S. Citizenship / No clearance needed / 100% remote within the US Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental US Type: Long-term contract (3+ years) Our client, a premier national healthcare provider, is currently...Long term contractFull timeRemote work
- Job Summary Job Summary The Support Lead (SRE) is responsible for overseeing the support operations and site reliability engineering tasks, ensuring the effective functioning of systems and applications. The primary goal is to enhance system performance, availability,...
$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...Currently hiringRemote work$62k - $141k
Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have...Full timeContract workPart timeWork at officeLocal areaRemote work$150k - $180k
...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business... ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or...Permanent employmentFull timeWork at officeLocal areaRemote workWorldwide$230k - $250k
GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is...Remote work$117.2k - $176.7k
...specific level of U.S. government background investigation and clearance required for this role.Overview of the Role:Join our Site Reliability Engineering (SRE) team, where you'll work alongside Infrastructure and Research & Development (R&D) partners to keep Salesforce...Full timeWork experience placement$190k - $225k
...Job TitleSRE + Release Pipeline EngineerJob DescriptionSRE + Release Pipeline Engineer Clearance: Top Secret clearance Location: Remote Role Framing Owns the path from local development to deployed AWS cluster across DCSA's GovCloud (IL2/IL5) and classified (IL6/Secret...Full timePart timeWork experience placementLocal areaRemote work- ...Infrastructure Reliability EngineerStack Infrastructure (Stack) provides digital infrastructure... ...for an Infrastructure Reliability Engineer with subject matter expertise in electrical... ...Development to enhance technical training for site teams based on lessons from event...Local areaFlexible hoursShift workNight shift
- ...Stanley Reid & Company is seeking a Systems Engineer (SRE) to support multiple cross-functional project teams in Chantilly, VA. You will monitor applications, coordinate changes to the environment, and drive resiliency and performance improvements in a cloud-native AWS...
$130k - $200k
Summary Position Title: Site Reliability Engineer Position ID: TA247 Location(s): On-site; Aurora, CO; Herndon, VA Application Deadline: August 31, 2026 Security Clearance Requirement: TS/SCI Security Clearance with Polygraph Job Description Trusted Space...Full timeTemporary workLocal area$86.8k - $198k
Release Train EngineerThe Opportunity:The Lead Release Train Engineer (RTE) serves as the Agile execution leader responsible for coordinating... ...our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen is...Full timeContract workPart timeWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site leader Fairfax, VA
- on-site clinical research associate (traveling/remote) Fairfax, VA
- official site Fairfax, VA
- historic site Fairfax, VA
- IT site lead Fairfax, VA
- junior website developer Fairfax, VA
- site safety Fairfax, VA
- site services specialist Fairfax, VA
- construction site safety Fairfax, VA
- site reliability engineer


