Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

ClearanceJobs

Senior Site Reliability EngineerThe Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability of cloud native systems. Our SREs don't just manage infrastructure; they build it using Infrastructure as Code (IaC), monitor it through advanced observability stacks, and protect it by engineering for failure. We work closely with clients to bridge the gap between development and operations. We are seeking a highly experienced and autonomous Senior Site Reliability Engineer (SRE) to join our dynamic team. As a technical leader, you will define the strategy and apply advanced software engineering principles to operations, focusing on the architecture, reliability, and long-term performance of large-scale production systems. You will play a crucial role in reducing toil through automation, defining and monitoring Service Level Objectives (SLOs), and implementing best practices for system stability and incident response. This role requires working with modern cloud technologies to ensure the high availability and efficiency of applications and infrastructure. Location: Primarily Remote. Candidates must be based in CA or DC Metro Area for proximity to project and client teams. Security Clearance Requirement: Applicants must be US citizens and eligible to obtain and maintain an active Secret security clearance or above.Key ResponsibilitiesSite Reliability EngineeringDrive the definition and adoption of SLIs and SLOs across multiple services or entire platforms, ensuring alignment with business goals.Design and architect Infrastructure as Code (IaC) solutions for large-scale, complex environments, establishing standards and best practices.Implement and manage containerized and serverless architectures using Docker, Kubernetes, and cloud-native services, focusing on performance and error budgets.Build and maintain reliable and self-healing CI/CD pipelines to automate deployments and improve development workflows.Toil Reduction and Incident ManagementImplement and refine comprehensive monitoring, alerting, and logging to detect and address performance and availability issues proactively.Lead the strategic effort to eliminate toil, identifying and championing major automation projects that deliver significant organizational efficiency.Lead high-severity incident response and coordinate blameless postmortems for major outages, driving the resulting remediation and systemic improvements.Testing and Service ResiliencyImplement cloud security best practices, including identity and access management (IAM), encryption, and compliance controls.Proactively identify and address system weaknesses and ensure performance under stress.Support disaster recovery and high availability strategies through backup and failover planning.Collaboration and Knowledge SharingServe as a primary SRE liaison for development teams, influencing application architecture and design to meet reliability and scalability targets from inception.Create and maintain documentation for cloud architectures, deployment processes, and best practices.Contribute to internal knowledge-sharing initiatives, ensuring continuous learning within the team.Stakeholder CommunicationAct as a subject matter expert and trusted advisor to clients and internal leadership on cloud infrastructure, reliability strategy, and Service Level Agreement (SLA) negotiations.Act on client feedback to refine and enhance cloud solutions.Conduct training and knowledge-sharing sessions to help clients manage their cloud environments effectively.Continuous Learning and InnovationStay updated on the latest developments in cloud infrastructure and technology trends.Drive innovation by proposing and implementing new techniques and technologies.Qualifications5+ years of experience in site reliability engineering, cloud engineering, or related fields.Strong software engineering skills with an emphasis on writing clean, modular, and maintainable code, specifically for automation and system management.Deep experience with Infrastructure as Code (IaC) tools like Terraform or CloudFormation.Deep experience with containerization and orchestration tools like Docker and Kubernetes.Deep knowledge of networking concepts, cloud security best practices, and identity management.Experience with programming or scripting languages such as Python, Bash, or Go.Experience with CI/CD pipelines and DevOps methodologies.Strong problem-solving skills and the ability to troubleshoot complex cloud environments.Demonstrated ability to influence technical decision-making across organizational boundariesPreferred qualifications:Bachelor's or advanced degree in Computer Science or a related field.Any of the below cloud certifications: Google Cloud Professional Cloud Architect Google Cloud Professional Cloud DevOps Engineer AWS Certified Solutions Architect AWS Certified Developer AWS Certified SysOps Administrator CompTIA Security+ certification or an equivalent DoD 8140/8570 IAT Level II baseline certification.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Fairfax, VA vacancy
  • $81.1k - $187k

     ...infrastructure and/or service according to terms for reliability and functionality.- Assists team members...  ...deployments.- Gains basic knowledge of site reliability trends and shares relevant...  ...are seeking a skilled Site Reliability Engineer to design, build, operate, and automate... 
    Suggested
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle Corporation

    Reston, VA
    3 days ago
  • $133k - $190k

    Site Reliability Engineer needed for a full time opportunity with SOC's direct client based in Herndon, VA. Direct Hire Role **Due to federal requirements, candidates must hold and possess an Active DOW TS/SCI security clearance to be considered for this role.** SOC is... 
    Suggested
    Full time

    SOC Support Services

    McLean, VA
    3 days ago
  • $119.8k - $234.7k

     ...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...opportunity for a Senior Site Reliability Engineer (SRE) to join the Azure Silver and Sovereign... 
    Suggested
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Reston, VA
    1 day ago
  • $135.8k - $183.8k

     ...dynamic and flexible work environment with competitive benefits and the ability to grow your career.We are looking for a Site Reliability Engineer to support our team responsible for building, managing, maintaining, deploying, and securing mission-critical services to... 
    Suggested
    Work at office
    Flexible hours

    VeriSign

    Reston, VA
    4 days ago
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and... 
    Suggested
    Permanent employment
    Full time
    Contract work
    Remote work
    Flexible hours

    Guidehouse

    McLean, VA
    2 days ago
  • We're seeking a skilled and proactive Site Reliability Engineer to join our team, ensuring the stability, security, and efficiency of our technological resources as we deliver cutting-edge AI solutions to the government. This is a fully remote position for candidates in... 
    Remote work

    Knexus

    Vienna, VA
    12 hours ago
  • $91.4k - $187k

    Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible... 
    Temporary work
    Flexible hours
    Shift work
    Weekend work

    Oracle Corporation

    Reston, VA
    12 hours ago
  • $146k - $194k

     ...focused on positioning Anduril as a lead provider of specialized engineering and products for Intelligence Community (IC) customers. We...  ...pressing national security requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure the health,... 
    Full time
    Work experience placement
    Immediate start
    Remote work

    Anduril Industries

    Reston, VA
    3 days ago
  • $109.18k - $163.77k

     ...channels. As a global company, we have offices in nine countries and can insert advertisements around the world.Job SummaryThe Site Reliability Engineering team is responsible for managing the critical infrastructure that powers FreeWheel's Streaming Hub platform. Streaming... 
    Full time

    Comcast

    Reston, VA
    12 hours ago
  • $84.24k - $142.48k

    OverviewJoin us to work collaboratively with our talented team of dynamic and passionate engineers to deliver capabilities that enable our customers to make a difference. You'll deploy and operate ArcGIS Velocity and ArcGIS Workflow Manager SaaS solutions. You will also... 
    Worldwide
    Flexible hours

    ESRI

    Vienna, VA
    12 hours ago
  • $87.1k - $157.45k

     ...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across...  ...to come in and help us build systems that stay reliable when things get complicated. We need a Site Reliability Engineer who has experience building, deploying... 
    Local area
    Immediate start
    Work from home
    Flexible hours

    Leidos

    Burke, VA
    3 days ago
  •  ...Senior Site Reliability EngineerThe Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long... 
    Remote work

    ClearanceJobs

    Fairfax, VA
    3 days ago
  • $158.5k - $230k

     ...extraordinary experiences together. Bring your whole self. The Role and Team We are growing our GovCloud team and looking for a Staff Site Reliability Engineer to help scale how we operate Medallia's US public-sector cloud platform. You will support federal agencies and other... 
    Permanent employment
    Temporary work
    Work experience placement
    Work at office
    Local area
    Remote work
    3 days per week

    Medallia

    McLean, VA
    3 days ago
  • $128.5k - $190k

     ...We empower exceptional people to create extraordinary experiences together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure and applications that power a highly reliable global SaaS... 
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    3 days ago
  • $121.5k - $264.1k

     ...sharing guidance on practices and terms for reliability and functionality.- Supervises team...  ...developing and maintaining knowledge of site reliability trends and sharing valuable...  ...Experience:9 years of experience in software engineering, infrastructure management, or related... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Reston, VA
    3 days ago
  • $102k - $234.6k

     ...members in designing and architecting infrastructure and service for reliability and functionality. Provides day-to-day direction to help...  ...to experiment with new technology, execute improvements, build site reliability knowledge, and provide clear data.Only Oracle brings... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Reston, VA
    1 day ago
  • $84.9k - $209.5k

     .... You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers...  ...posted.Career Level - IC4Escalation points for junior site reliability engineers during complex or high-impact incidents.Manage and... 
    Temporary work
    Monday to Friday
    Flexible hours
    Shift work
    Night shift

    Oracle Corporation

    Reston, VA
    1 day ago
  • $128.83k - $193.25k

     ...you will be responsible for ensuring the reliability, scalability, and performance of our data systems. Working closely with data engineers and other operation sub-teams, you will manage...  ...and benefits summary on our careers site for more details.EducationBachelor's DegreeWhile... 
    Full time

    Comcast

    Reston, VA
    12 hours ago
  •  ...Required U.S. Citizenship / No clearance needed / 100% remote within the US  Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental US  Type: Long-term contract (3+ years) Our client, a premier national healthcare provider, is currently... 
    Long term contract
    Full time
    Remote work

    ASCENDING

    Fairfax, VA
    6 days ago
  • Job Summary Job Summary The Support Lead (SRE) is responsible for overseeing the support operations and site reliability engineering tasks, ensuring the effective functioning of systems and applications. The primary goal is to enhance system performance, availability,... 

    TechDigital Group

    Fairfax, VA
    1 day ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    1 day ago
  • $62k - $141k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Chantilly, Loudoun County, VA
    4 days ago
  • $150k - $180k

     ...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business...  ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    12 hours ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Remote work

    Govcio

    Arlington, VA
    12 hours ago
  • $117.2k - $176.7k

     ...specific level of U.S. government background investigation and clearance required for this role.Overview of the Role:Join our Site Reliability Engineering (SRE) team, where you'll work alongside Infrastructure and Research & Development (R&D) partners to keep Salesforce... 
    Full time
    Work experience placement

    Salesforce

    Herndon, VA
    12 hours ago
  • $190k - $225k

     ...Job TitleSRE + Release Pipeline EngineerJob DescriptionSRE + Release Pipeline Engineer Clearance: Top Secret clearance Location: Remote Role Framing Owns the path from local development to deployed AWS cluster across DCSA's GovCloud (IL2/IL5) and classified (IL6/Secret... 
    Full time
    Part time
    Work experience placement
    Local area
    Remote work

    ClearanceJobs

    Arlington, VA
    3 days ago
  •  ...Infrastructure Reliability EngineerStack Infrastructure (Stack) provides digital infrastructure...  ...for an Infrastructure Reliability Engineer with subject matter expertise in electrical...  ...Development to enhance technical training for site teams based on lessons from event... 
    Local area
    Flexible hours
    Shift work
    Night shift

    STACK Infrastructure

    Manassas Park, VA
    1 day ago
  •  ...Stanley Reid & Company is seeking a Systems Engineer (SRE) to support multiple cross-functional project teams in Chantilly, VA. You will monitor applications, coordinate changes to the environment, and drive resiliency and performance improvements in a cloud-native AWS... 

    Stanley Reid & Company

    Chantilly, Loudoun County, VA
    12 hours ago
  • $130k - $200k

    Summary Position Title: Site Reliability Engineer Position ID: TA247 Location(s): On-site; Aurora, CO; Herndon, VA Application Deadline: August 31, 2026 Security Clearance Requirement: TS/SCI Security Clearance with Polygraph Job Description Trusted Space... 
    Full time
    Temporary work
    Local area

    Trusted Space, LLC

    Herndon, VA
    12 hours ago
  • $86.8k - $198k

    Release Train EngineerThe Opportunity:The Lead Release Train Engineer (RTE) serves as the Agile execution leader responsible for coordinating...  ...our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen is... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Herndon, VA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!