Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$190k - $225k

ClearanceJobs

SRE + Release Pipeline EngineerClearance: Top Secret clearance Location: Remote Role Framing Owns the path from local development to deployed AWS cluster across DCSA's GovCloud (IL2/IL5) and classified (IL6/Secret) partitions, and the health of those clusters. These environments are currently in IATT and working towards scale and ATO. Deliberately a single role: at current team size, build/release and run/operate are the same person's problem. Release side: CDK stacks, Flux GitOps, Helm chart authoring, Envoy Gateway routes, EKS/IRSA, container builds, GitLab CI pipeline automation, and the local-to-AWS promotion path. Run side: the observability/tracing stack, cross-service tracing, runtime health, credential rotation, incident response, and the preflight/diagnose/QA loops. This role directly supports OY2 Workstream 1 (Image Builder Pipeline — STIG-compliant AMI automation, Artifactory integration, GitLab CI deployment), Workstream 6 (Import Account Standardization — Baseline Account Pipeline, tagging compliance, centralized VPC endpoints), and the GenAI IDP deployment pipeline (deploying the IDP Solution to Customer non-production and production environments via IaC).Basic Qualifications 3+ years of professional experience in SRE, DevOps, platform, or infrastructure engineering Experience operating Kubernetes (EKS) in production, including in DoD/IC classified environments Experience packaging and deploying applications with Helm (authoring and maintaining charts, not only consuming them) Experience with Flux (or an equivalent GitOps controller — e.g., Argo CD) driving continuous delivery of Helm releases Experience with AWS compute and managed services (e.g., EKS, RDS, S3, IAM/IRSA, EC2 Image Builder) Experience with infrastructure-as-code; TypeScript/CDK experience specifically, or demonstrated ability to work in a TypeScript IaC codebase Experience with production observability and distributed tracing (e.g., OpenTelemetry, Grafana/Tempo, or equivalent), used to diagnose failures from telemetry rather than guesswork Experience leading incident diagnosis and resolution, including identifying and confirming root cause before remediating Experience with GitLab CI/CD pipelines for build automation and deployment Proficiency in at least one scripting language (e.g., Python) for tooling and automationPreferred Qualifications Experience building STIG-compliant AMI pipelines using EC2 Image Builder with DoD security baseline validation Experience with Artifactory integration for AMI/container image distribution across AWS Organizations Experience with cross-domain solutions (AWS Diode, CDS) and SIPRNet/JWICS environments Experience with AWS Organizations, SCPs, OU design, and multi-account governance Experience with certificate lifecycle management (ACM, Private CA) and automated renewal workflows Experience deploying or operating ML/LLM-serving infrastructure (e.g., Amazon Bedrock endpoints, SageMaker, model-inference endpoints) and reasoning about token/latency/cost from traces Experience with an ECR- or registry-based GitOps model (charts/images mirrored to a registry, controller reconciling from it) Experience diagnosing failed or stuck Helm/Flux reconciliations (HelmRelease not converging, drift between desired and live state) Experience with Envoy/gateway and certificate/TLS management in classified environments Proficiency using AI coding assistants as a daily driver, with the judgment to validate their output before it reaches a cluster Bachelor's degree in Computer Science or a related field, or equivalent practical experienceProgram Qualification Run an engineering practice driven by evaluation, testing, and verification — changes are proven with tests, traces, and metrics before they are called done Operate agentic systems fluently (MCP tools, Bedrock/LLM calls, streaming, distributed traces) in support of GenAI IDP document classification and field extraction workflows Work within DCSA classified environments (IL5/IL6) following DoD security baselines and STIG compliance requirements Support the GenAI IDP solution: prompt engineering, model fine-tuning, evaluation logic for document section classification and extracted field validation Use AI coding assistants heavily as daily drivers, while remaining skeptical of their output and holding it to the same evidentiary bar as any other code Current, active Top Secret security clearance Current, active Security+ or equivalent certificate for privileged user access Compensation and Benefits Salary Range: $190,000 - $225,000 (Compensation is determined by various factors, including but not limited to location, work experience, skills, education, certifications, seniority, and business needs. This range may be modified in the future.) Benefits: Gridiron offers a comprehensive benefits package including medical, dental, vision insurance, HSA, FSA, 401(k), disability & ADD insurance, life and pet insurance to eligible employees. Full-time and part-time employees working at least 30 hours per week on a regular basis are eligible to participate in Gridiron's benefits programs. Gridiron IT Solutions is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status or disability status. Gridiron IT is a Women Owned Small Business (WOSB) headquartered in the Washington, D.C. area that supports our clients' missions throughout the United States. Gridiron IT specializes in providing comprehensive IT services tailored to meet the needs of federal agencies. Our capabilities include IT Infrastructure & Cloud Services, Cyber Security, Software Integration & Development, Data Solution & AI, and Enterprise Applications. These capabilities are backed by Gridiron IT's experienced workforce and our commitment to ensuring we meet and exceed our clients' expectations.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Arlington, VA vacancy
  • $115.5k - $164.8k

     ...mission that matters at a company where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant...  ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    1 day ago
  • $133k - $190k

    Site Reliability Engineer needed for a full time opportunity with SOC's direct client based in Herndon, VA. Direct Hire Role **Due to federal requirements, candidates must hold and possess an Active DOW TS/SCI security clearance to be considered for this role.** SOC is... 
    Suggested
    Full time

    SOC Support Services

    McLean, VA
    2 days ago
  • $165k - $230k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    1 day ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Suggested
    Remote work

    Govcio

    Arlington, VA
    4 days ago
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and... 
    Suggested
    Permanent employment
    Full time
    Contract work
    Remote work
    Flexible hours

    Guidehouse

    McLean, VA
    1 day ago
  • $185k - $230k

    As a Sr. Site Reliability Engineer (SRE) III, you’ll work as part of a collaborative and high-performing team providing your expertise to deliver technical solutions within the highest levels of the federal government.We know that you can’t have great technology services... 
    Full time
    Local area
    Immediate start

    MetroStar Systems

    Washington DC
    4 days ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    4 days ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    3 days ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Temporary work

    Kong

    Washington DC
    3 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    12 hours ago
  • $150k - $180k

     ...Umbra.About the JobWe are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business...  ...impact across the organization.This position is based on-site in either our Arlington, VA office, Reston, VA office or... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Worldwide

    Umbra

    Arlington, VA
    3 days ago
  • $125k - $185k

     ...lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    4 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions...  ...locate missing children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  •  ...A leading security infrastructure firm in Washington, D.C. is seeking a hands-on Site Reliability Engineer (SRE) with expertise in Kubernetes and cloud infrastructure. The role emphasizes total ownership of security infrastructure while defending against advanced threats... 

    Cyrad Solutions LLC

    Washington DC
    4 days ago
  • $100k - $110k

     ...for new hire onboarding and occasional in-person team meetings and company events. We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will... 
    Permanent employment
    Remote work
    Flexible hours

    GrabJobs

    Washington DC
    5 days ago
  • $197.3k - $225.1k

     ...through technology, we equally prioritize cybersecurity, reliability, software quality, and data management. Technology & Data...  ...are highly-skilled information security, cybersecurity, site reliability engineering, technology, data analyst, data scientist, and risk management... 
    Full time
    Part time
    Local area

    Socket

    McLean, VA
    4 days ago
  • $158.5k - $230k

     ...together. Bring your whole self. The Role and Team We are growing our GovCloud team and looking for a Staff Site Reliability Engineer to help scale how we operate Medallia's US public-sector cloud platform. You will support federal agencies and other... 
    Permanent employment
    Temporary work
    Work experience placement
    Work at office
    Local area
    Remote work
    3 days per week

    Medallia

    McLean, VA
    2 days ago
  • $82.3k - $228.8k

     ...inclusive environment, empowering our employees to be their authentic selves. We are seeking a highly experienced Senior Site Reliability Engineer – Compute Platforms to design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud... 
    Temporary work
    Work at office
    Remote work
    Worldwide
    3 days per week

    GrabJobs

    Washington DC
    1 day ago
  •  ...certificates. Qualifications & Requirements ~ Bachelor’s degree in Computer Science, Engineering, or a related technical field. ~3+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or infrastructure-focused roles. ~ Hands-on... 
    Temporary work

    2T Consulting

    Washington DC
    6 days ago
  • $87.1k - $157.45k

     ...throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across...  ...to come in and help us build systems that stay reliable when things get complicated. We need a Site Reliability Engineer who has experience building, deploying... 
    Local area
    Immediate start
    Work from home
    Flexible hours

    Leidos

    Falls Church, VA
    2 days ago
  • $128.5k - $190k

     ...exceptional people to create extraordinary experiences together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure and applications that power a highly reliable global SaaS... 
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    1 day ago
  • $145k - $175k

     ...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data... 
    Full time
    Remote work

    GrabJobs

    Washington DC
    5 days ago
  • $175k - $195k

     ...products need key features. They need to be reliable, scalable, performant, cost effective,...  ...for thinking through these problems and engineering solutions to them. We hire excellent...  ...completed or currently under development.Site Reliability EngineerAs a Site Reliability... 
    Full time
    Temporary work
    Work experience placement
    Remote work

    Filevine

    Washington DC
    5 days ago
  •  ...bottlenecks, and improve system health—utilization, performance, and reliability—across our infrastructure Understand how systems fail and...  ..., documentation, and code review—that automates reliability engineering work: Deployment tooling Fault-injection/chaos... 
    Full time
    Casual work
    Local area
    Worldwide
    Flexible hours
    Shift work

    Ping Identity

    Washington DC
    11 days ago
  •  ...Job Description Job Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient... 
    Local area

    Tiger Analytics Inc.

    Washington DC
    more than 2 months ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    15 days ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    28 days ago
  • $207k - $284.9k

     ...on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity is...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government customers... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    1 day ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  • $166k - $220k

     ...failure. As such, it is critical that Anduril services are reliable and maintainable. This means that all services &...  ...ground systems & Kubernetes infrastructure.ABOUT THE JOBAs a Site Reliability Engineer on the Observability team, you will build & operate Anduril... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!