Site Reliability Engineer
$180k - $270kStrategic Business Systems
AWS DevOps / Agentic SRE Engineer (Active TS/SCI Clearance)
- Chantilly, Virginia
AWS DevOps / Agentic SRE Engineer – TS/SCI
Clearance: Active TS/SCI
Certification: Active Security+ or equivalent required
Experience: 3+ years of SRE, DevOps, Platform Engineering, or Infrastructure Engineering experience
Position Overview
We are seeking an SRE / DevOps & Release Engineer to own the deployment, reliability, observability, and operation of secure AWS environments supporting an advanced Agentic AI platform.
This is a hands‑on engineering position combining DevOps, Site Reliability Engineering, platform engineering, and release automation. The engineer will own the path from local development through production deployment into AWS Kubernetes environments, as well as the operational health and reliability of those environments.
The platform is currently operating within IATT environments and progressing toward scale and ATO. At the current team size, build/release and production operations are intentionally combined — the engineer responsible for deploying the platform also has responsibility for ensuring that it operates reliably.
The environment includes AWS EKS, CDK, Kubernetes, Helm, Flux GitOps, container registries, Envoy Gateway, distributed tracing, and modern observability technologies.
Key Responsibilities
- Design, build, maintain, and operate highly available AWS infrastructure supporting an Agentic AI platform.
- Own the software delivery lifecycle from local development through build, test, packaging, promotion, and deployment into AWS environments.
- Deploy and operate Kubernetes workloads in production AWS EKS environments.
- Author, maintain, and troubleshoot Helm charts supporting platform applications and services.
- Build and maintain GitOps-based continuous delivery utilizing Flux, Helm, and container registries.
- Develop and maintain Infrastructure as Code using AWS CDK and TypeScript.
- Build and manage AWS infrastructure utilizing EKS, RDS, S3, IAM/IRSA, ECR, and related AWS services.
- Build and maintain container-image and Helm-chart promotion processes.
- Configure and support Envoy Gateway routing, certificates, and TLS.
- Implement and operate observability and distributed-tracing capabilities utilizing OpenTelemetry, Grafana, Tempo, or equivalent technologies.
- Establish cross-service tracing to provide end‑to‑end visibility into distributed application and agent workflows.
- Monitor cluster, application, and service health and proactively identify reliability issues.
- Diagnose failed or stuck Helm and Flux reconciliations, configuration drift, pod failures, and deployment issues.
- Lead production incident investigation and root‑cause analysis using Kubernetes state, logs, metrics, and distributed traces.
- Develop preflight validation, diagnostic, QA, and operational tooling.
- Automate infrastructure and operational processes using Python and other scripting technologies.
- Support credential, certificate, and secrets rotation.
- Partner with software engineers and AI/ML teams to deploy and operate agentic applications, model‑serving infrastructure, and supporting services.
- Support platform security, compliance, IATT, and ATO activities.
- Use AI‑assisted development tools to accelerate engineering while independently validating generated code and configuration before deployment.
- Drive an engineering approach based on measurable evidence, automated testing, telemetry, and verification.
Required Qualifications
- Current and active TS/SCI security clearance
- Current Security+ certification or equivalent certification supporting privileged‑user access.
- 3+ years of professional experience in Site Reliability Engineering, DevOps, Platform Engineering, Cloud Engineering, or Infrastructure Engineering.
- Hands‑on experience operating Kubernetes in production environments.
- Experience authoring and maintaining Helm charts.
- Experience implementing GitOps‑based continuous delivery using Flux, Argo CD, or equivalent technologies.
- Hands‑on AWS experience with services such as EKS, RDS, S3, IAM/IRSA, and ECR.
- Experience developing and maintaining Infrastructure as Code.
- Experience working with TypeScript/AWS CDK or demonstrated ability to work within a TypeScript‑based IaC environment.
- Experience implementing and utilizing production observability and distributed‑tracing technologies such as OpenTelemetry, Grafana, and Tempo.
- Demonstrated experience diagnosing infrastructure and application failures using telemetry, logs, metrics, and traces.
- Experience leading incident investigation, root‑cause analysis, remediation, and validation.
- Proficiency with Python or another scripting language for infrastructure automation and operational tooling.
- Strong understanding of CI/CD, containers, networking, security, and modern cloud architecture.
Preferred Qualifications
- TS/SCI with Poly
- Experience implementing registry‑based GitOps architectures utilizing ECR or similar container registries.
- Experience troubleshooting Flux and Helm reconciliation issues, including HelmRelease failures and configuration drift.
- Experience deploying and operating ML/LLM‑serving infrastructure such as KServe, MLflow, or model‑inference endpoints.
- Experience supporting AI or Agentic AI platforms.
- Familiarity with Amazon Bedrock, LLM services, MCP tools, and agent‑based architectures.
- Ability to analyze model and agent performance characteristics such as latency, token utilization, and cost using distributed traces.
- Experience with Envoy, API gateways, routing, certificate management, and TLS.
- Experience supporting secure Federal or Intelligence Community AWS environments.
- Experience supporting systems through IATT and ATO processes.
- Experience using AI coding assistants while independently validating generated code and infrastructure changes before deployment.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical experience.
Why SBS
SBS is a trusted AWS partner supporting both federal and commercial customers. Our teams focus on delivery—building and operating secure, scalable solutions that matter.
You’ll be part of a team that:
- Works directly with AWS on meaningful, mission‑driven programs
- Operates in high‑impact national security environments
- Builds and operates modern AWS and Kubernetes platforms
- Works across cloud infrastructure, DevOps, SRE, GitOps, and AI
- Values engineers who can build, automate, troubleshoot, and deliver
COMPENSATION & BENEFITS
SBS offers a comprehensive total‑rewards package, including a market competitive salary along with:
Comprehensive medical, dental, and vision coverage; HSA‑eligible plan options available
401(k) retirement plan with company match (vesting schedule per Plan Document)
Paid Time Off, federal holidays, and floating holiday for personal observance
Annual professional development support for AWS certifications, training, and conferences
Employee referral program where applicable and documented by program policy
Life, AD&D, and short- / long‑term disability insurance
Telework and flexible‑schedule support where mission and contract permit
Mission‑focused federal contractor supporting national‑security customers
Target Salary Range: $180,000-270,000. This range reflects the anticipated compensation for this role. Actual salary will be based on a combination of factors, including the position’s scope and level of responsibility, the candidate’s relevant experience, education, technical expertise, skills and qualifications, geographic location, and applicable business or contractual requirements.
About SBS
Strategic Business Systems, Inc. (SBS) is a national Information Technology services company headquartered in the Washington, D.C. metropolitan area. SBS provides IT infrastructure design, integration, and operational services. Our expertise spans the full spectrum of infrastructure technologies, including networking, servers, data storage, disaster recovery, cybersecurity, and internet technologies.
Equal Employment Opportunity
SBS is an equal opportunity employer; all qualified applicants will receive consideration for employment without regard to age, gender, gender identity, sex, sexual orientation, color, race, creed, national origin, religion, marital status, parental status, citizenship status, ancestry, physical or mental disability, genetic information, veteran status, military status, or any other classification protected by federal, state, or local laws.
Accommodations
If you need an accommodation while seeking employment with SBS, please email View email address on click.appcast.io . Accommodations are made on a case‑by‑case basis.
No Unsolicited Agency Referrals
SBS does not accept unsolicited resumes from staffing agencies. Any resumes submitted without a prior agreement will be considered the property of SBS.
$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...such as Python, Go, Perl, or Ruby. ~ Mentored other engineers and influenced technical direction through design reviews, architecture...Suggested- ...Early Warning Services LLC in Scottsdale seeks a Principal Site Reliability Engineer to design high-availability systems and scale microservice architectures. You will partner with development teams to implement observability, automation, and resilient deployment patterns...Suggested
$182.8k - $247.3k
...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...SuggestedWork experience placement$191k - $253k
...TEAM: CorpTech Platform is the internal engineering force multiplier behind Anduril’s... ...THE JOB: This Staff SRE role sets the reliability architecture for the systems that run Anduril... ...: ~10+ years of experience in site reliability engineering, production engineering...SuggestedFull timeWork experience placementImmediate start$194k - $237k
## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on: Posted 5 Days Agojob requisition id: REQ2026... ...sponsorship.**Overall Purpose**The Principal Site Reliability Engineer partners with development teams by designing availability and...SuggestedHourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$194k - $267k
...Staff Site Reliability Engineer - Kubernetes Important: if an employer asks you to log into their system via iCloud or Google, send a code, an SMS or Telegram password, run some code, or install software — refuse. These are signs of fraud. Secure Every Identity,...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$181k - $265k
...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to... ...and Helm Understand the importance of performant and reliable systems Education - Ideally looking for a B.A. / B.S. degree...Work at officeImmediate start3 days per week- ...NVIDIA Corporation is seeking a Sr. Systems Software Engineer to build platform software around open-source container runtimes and Kubernetes. You will contribute to GPU-accelerated applications, work on Kubernetes integration, and collaborate across NVIDIA teams to...
- ...NVIDIA Corporation is seeking a Senior System Software Engineer for the Autonomous Vehicles Platform in Santa Clara, CA. You will develop high-performance, scalable system software for data collection and fleets of autonomous vehicles, interfacing with sensors and heterogeneous...
$136k - $218.5k
## Systems Quality and Reliability Engineer - LPUApplylocations: US, CA, Santa Claratime type: Full timeposted on: Posted Todayjob requisition id: JR2018660We are seeking Systems Quality and Reliability Engineer to join our LPU team!NVIDIA has continuously reinvented itself...- ...AZX in Seattle is seeking a Software Engineer to build models and implement an agentic, formally-grounded solution that addresses client... ..., packaging, testing, CI, and documentation, delivering a reliable interface for AZX engineers to provide provably correct answers...
$185k - $260k
...LinkedIn Top Startups. Your Role As a Senior Software Engineer on the Developer Platform team, you will be responsible for designing... ...initiatives to ensure that new features are performant and reliable, creating innovative frameworks to address our unique testing...Immediate startRemote workHome officeFlexible hours- ...Vercel is seeking a Staff level Engineer to lead the Finance Data Platform, powering revenue reporting and compliance across the business. Define the technical vision and build auditable pipelines with Kafka, ClickHouse, Tinybird, and Snowflake. You will mentor engineers...
- ...Trustly, Inc. is looking for a Staff AI Enablement Engineer to shape how teams across the company adopt AI at scale. You will design internal... ..., mentoring ability, and a passion for turning prototypes into reliable production systems. You’ll partner with product and engineering...
- ...Sprig is hiring for a Platform Build Engineer to lead the next era of fast, reliable CI/CD and hermetic builds for a massive TypeScript/React/Go monorepo. You will reduce build times, make tests reliable, and create ephemeral environments per task to accelerate developer...
- ...Notion is seeking a Developer Platform engineer to build tools, APIs, and platform experiences that connect Notion to the world — bringing... ...work will enable developers, admins, and builders to create reliable integrations and automations on top of Notion. You’ll contribute...
- ...EvenUp Inc is hiring a Senior Platform Foundation Engineer to lead the architecture of core platform capabilities powering the whole system. You will own the stack from design through deployment, enabling other teams with APIs, data exports, and streaming capabilities...
$200k - $250k
...fastest-growing tech companies, bringing together early-stage engineers, product builders, and business athletes from companies like Stripe... ...meetings, and Gigs Republic, our bi-annual company off-site . Our offices are designed to feel like home-inspired workspaces...Work at officeRemote workWork from homeRelocationHome office$228.5k
...you see yourself as a mission-driven Senior Manager, Platform Engineering who believes that great products emerge from strong teams? Do... ...success. We build complex, data-driven systems that need to work reliably during critical moments, scale as our partners grow, and...Full timeTemporary workRemote workHome officeFlexible hoursShift work- ...Yahoo Inc. is seeking a Partner Success Engineer to own end-to-end partner integrations for Yahoo Mail’s data platform, ensuring API... ...collaboration with enterprise teams and executives. You will drive reliability, specify data contracts, and mentor teams while shaping...
$200k - $240k
...is today’s application monitoring standard and our team is building its AI-native future. About the role: The Solutions Engineering team at Sentry is responsible for helping our largest customers successfully embed Sentry into their applications and workflows,...Hourly payFlexible hours- ...Cencora seeks a Forward Deployed Engineer to embed with business units and deliver production-grade Generative AI solutions in days to weeks... ...engineering, solution architecture, and consulting, with on-site engagement as needed and a strong emphasis on delivering value quickly...
- ...Databricks is hiring an AI Forward Deployed Engineer (FDE) to design, build, and productionize GenAI applications for the U.S. federal sector... ...Washington, D.C./Maryland/Virginia metro area with periodic on-site work. Remote work is possible. You will own end-to-end GenAI...Remote work
$191k - $253k
...vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE TEAM At Anduril, our Software Engineers specialize in solving complex, real-world problems through cutting-edge algorithms and intelligent software integrations....Full timeWork experience placementImmediate start$200k - $250k
...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to... ...it out on their hardest use case, and get agents running reliably at scale. This is a hands-on, highly technical team. Solutions...Work at officeFlexible hours- ...team combines deep expertise in model innovation and systems engineering paired with a design-minded product engineering team to build... ...of the business while maintaining a high bar for performance, reliability, and engineering quality. What You Bring Given the...Work at officeVisa sponsorshipFlexible hours
- ...journey. About the Role We're hiring a Senior Software Engineer to join Pivotal's Case Execution team. This team owns critical... ...executed end to end, with a clear eye toward correctness and reliability. Collaborate cross-functionally: Work closely with the Integrations...Remote workFlexible hours
$186.07k - $218.9k
...surges.” learn more about working at Coinbase. Senior Software Engineer (EAA) The EAA Compliance CXAE team, part of Coinbase's... ...capabilities that benefit multiple teams across the organization. Own reliability for Tier‑1 compliance systems by anticipating potential issues...Local area- ...Material Security, Inc. in San Francisco, CA seeks a Senior Software Engineer II to design and maintain scalable, cloud-based systems. You... ...and Agile practices, and collaborate across teams to deliver reliable, scalable software. 3+ years of core experience are required...
- ...the Team The Private Computing team works across product, engineering, security, and safety to build advanced privacy products and... ...integrity infrastructure Operate systems at scale with high reliability, including an on‑call rotation Collaborate with a diverse...Work at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- construction site safety Eastern, KY
- on-site clinical research associate (traveling/remote) Eastern, KY
- historic site Eastern, KY
- official site Eastern, KY
- site safety Eastern, KY
- junior site reliability engineer
- site reliability engineer
- site reliability engineering manager
- site reliability engineer remote
- lead site reliability engineer

