Site Reliability Engineer
EXIGER
Who We Are:
Exiger transforms supply chains into a strategic advantage, advancing our mission to make the world a safer and more transparent place to succeed. OurAI platform, 1Exiger, delivers instant visibility into complex supplier ecosystems, leveraging proprietary data and advanced AI to surface risk, automate compliance, and unlockefficiencies and cost savings to strengthen long-term resilience. Trusted by 550+ global customers, including Fortune 500 companies and U.S. government agencies, Exiger is a recognized, award-winning leader in supply chain AI and a FedRAMP authorized provider to the federal government.
Site Reliability Engineer
Location: U.S. (Hybrid)
This role requires U.S. citizenship and eligibility for a U.S. security clearance.
Role Summary:
Exiger is transforming how governments and global enterprises manage supply chain, defense, and geopolitical risk. Our AI-powered platform equips the world's most important institutions with the intelligence they need to protect critical infrastructure, secure national interests, and make data-driven operational decisions.
From identifying counterfeit parts in defense supply chains to anticipating geopolitical risk exposure, Exiger enables mission owners to act with clarity and confidence in complex, high-stakes environments.
This is our first dedicated Site Reliability Engineering hire and a founding role. You will help stand up the SRE function at Exiger: setting the standards, tooling, and practices that keep 1Exiger reliable for our 550+ customers, including Fortune 500 companies and U.S. government agencies. You will own reliability across the full service lifecycle, from design and capacity planning through deployment, monitoring, and incident response, and build the automation that lets the platform scale without scaling headcount. Because you are first, we need someone who has practiced SRE before and can bring the playbook, not learn it on the job.
You will use your expertise in coding, algorithms, complexity analysis, and large-scale distributed system design to solve the reliability challenges that are unique to operating a mission-critical AI platform in regulated and government environments.
SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.
What You'll Do:
- Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards that other engineering teams adopt.
- Build and own observability: instrument services for availability, latency, and system health, and turn that signal into actionable insight.
- Drive decisions with data: form hypotheses, measure the impact of every change, and let metrics rather than intuition set reliability priorities.
- Own the reliability of production services from design consulting and launch reviews through steady-state operation.
- Eliminate repetitive manual operations through automation and infrastructure as code, replacing them with reliable, self-service tooling.
- Plan for scale: capacity planning, performance analysis, and driving changes that improve both reliability and delivery velocity.
- Improve resilience through chaos engineering and fault-injection testing, running game days that prove the platform degrades gracefully and recovers from failure.
- Lead sustainable, blameless incident response and postmortems, and stand up and participate in an on-call rotation.
- Leverage AI-assisted development tooling (such as Codex and Claude) to accelerate automation, tooling, and investigation work, and help the team adopt these tools effectively.
What You Need:
- Bachelor's or Master's degree in Computer Science, a related field, or equivalent practical experience.
- 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role. As our first SRE hire, you must have practiced SRE before and be ready to establish the function.
- 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems.
- Strong grounding in Unix/Linux internals (filesystems, processes, system calls) and networking fundamentals (TCP/IP, DNS, routing, load balancing).
- Hands-on experience establishing core SRE practices from the ground up: SLIs, SLOs, and error budgets, monitoring and observability, capacity planning, and automation that removes repetitive manual work.
- A rigorous, empirical mindset: you form hypotheses, measure outcomes, and make metrics-driven decisions rather than relying on intuition or anecdote.
- Experience with chaos engineering or fault-injection testing (for example game days, Chaos Monkey, Gremlin, or LitmusChaos) to validate system resilience.
- Proven incident management experience: on-call ownership, leading response under pressure, and driving blameless postmortems to root cause.
- Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems.
- Familiarity with cloud platforms (AWS) and secure system integration.
- Comfort integrating AI coding assistants (such as Claude and Codex) into your daily engineering workflow.
- Ability to translate ambiguous mission problems into structured technical solutions.
- Ability to operate independently in dynamic, high-stakes environments.
- Willingness to travel as needed to support customer engagements.
Nice to Have:
- 4 years of experience programming in Go or C (Java also welcome), with the ability to debug, optimize, and automate rather than just script.
- Experience supporting ML or data platforms in production.
- Familiarity with data warehouses such as Snowflake, Redshift and/or Apache Iceberg.
- Experience operating in FedRAMP or other regulated or government environments.
Why You'll Love Working at Exiger:
- High-performance culture rooted in accountability, collaboration, and a shared commitment to excellence.
- Discretionary Time Off for all employees, with no maximum limits on time off
- Industry leading health, vision, and dental benefits
- Competitive compensation package
- 16 weeks of fully paid parental leave
- Flexible, hybrid approach to working from home and in the office where applicable
- Focus on wellness and employee health through stipends and dedicated wellness programming
- Purposeful career development programs with reimbursement provided for educational certifications
#LI-hybrid
Exiger is named a Leader in the GartnerMagic Quadrant for Supplier Risk Management, twice selected as one of Fast Company's 'Brands That Matter,' and recipient of the Third Party Risk Association's Innovator Award, Exiger's technology has been recognized by leading analyst evaluations and 50+ awards. Learn more at Exiger.com and follow Exiger on LinkedIn .
At Exiger, our values define how we work and why we lead. We are mission-inspired, imagination-driven, trust-anchored, and compassion-focused-committed to building technology that makes the world safer, more transparent, and more resilient.
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law.
Exiger's hybrid work policy is periodically reviewed and adjusted to align with evolving business needs.
- ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running... ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident...SuggestedFull timeWork at officeRemote workFlexible hours
- ...Site Reliability Engineer (SRE) Remote No sponsorship available. Must be able to obtain a Public Trust clearance. What You Will Do We are seeking a Site Reliability Engineer (SRE) to support the SBA Disaster Lending Platform modernization effort in a remote...SuggestedLocal areaRemote work
$81.1k - $187k
...infrastructure and/or service according to terms for reliability and functionality. - Assists team... .... - Gains basic knowledge of site reliability trends and shares relevant information... ...are seeking a skilled Site Reliability Engineer to design, build, operate, and automate...SuggestedTemporary workImmediate startFlexible hoursShift work$114.6k - $190.2k
...with MANTECH! ***This is for a future opportunity*** MANTECHseeks motivated, career, and customer-oriented Site Reliability Engineer (SRE) for a new initiative. This effort supports the rapid design, deployment, operation, and sustainment of enterprise-...SuggestedHourly payContract workTemporary workWork experience placementWork at officeLocal areaRemote work- Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments and predicting.Suggested
- ...grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers...Work experience placementRemote workFlexible hours
- ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment...Work experience placement
- ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington D.C. area, focusing on optimizing the availability, performance, and scalability of critical production services. The ideal...
$165k - $241.4k
...can only be performed by a U.S. citizen on U.S. soil. From a reliability standpoint, this role involves evaluating the scalability,... ...coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Permanent employmentFull timeTemporary workLocal areaFlexible hours- ...About the job Site Reliability Engineer (SRE) ***W2 only*** Position: Site Reliability Engineer (SRE) Work Authorization: All Work Authorizations Location: Reston, VA Contract: 24 months Description: Site Reliability Engineer (SRE) roles and...Contract work
$175k - $250k
Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or... ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design,...Full timeRemote workRelocationRelocation package- ...support Pension plan Paid maternity leave 401(k) Get notified when a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle, WA $115,000.00-$175,000.00 5 months ago Senior ServiceNow...Contract workRemote work
$90k - $150k
...Top Workplaces honoree, is seeking a SRE Engineer to support our growing team. The SRE... ...This role is responsible for improving the reliability, availability, performance,... ...Government customers. Work Environment: On-site Key Responsibilities: Define,...Permanent employmentFull timeContract work$51.9 per hour
...OVERVIEW: This job is responsible for the reliability, availability, and performance of... ...operational efficiency. This role blends software engineering, clinical engineering, and security... .... Works cross-functionally with AHN site leaders and teams to navigate and to monitor...For contractorsLocal area$178k - $213k
...Splunk Ventures, and Vista Credit Partners of Vista Equity Partners 2022 Cybersecurity Excellence Award for MDR Manager, Site Reliability Engineering Reports to: VP, Product Engineering Location: While proximity to Tampa is preferred to support hybrid schedule in Tampa...Permanent employmentWork experience placementWork at officeRemote workWork from homeHome officeFlexible hours- ...architecting infrastructure and service for reliability and functionality. Provides day-to-day... ...technology, execute improvements, build site reliability knowledge, and provide clear... ...: 8 years of experience in software engineering, infrastructure management, or related fieldORBachelor...Immediate startFlexible hours
$220k - $250k
...Staff Site Reliability Engineer Yugabyte is the company behind YugabyteDB, the AI-ready, multi-modal, distributed PostgreSQL database for cloud-native apps. Trusted by industry leaders including Shopify, Paramount+, GM, Kroger, Fiserv, and NPCI, YugabyteDB has been...H1bLocal areaWorldwideVisa sponsorship- ...TENEX Staff Site Reliability Engineer TENEX is an AI-native, automation-first, built-for-scale Managed Detection and Response (MDR) provider. We are a force multiplier for defenders, helping organizations enhance their cybersecurity posture through advanced threat detection...Work from home
- ...Lead Site Reliability Engineer Bridge Defense is redefining how modern defense technology is delivered. Based in Washington, D.C., we are built for the dynamic mission environment facing the Department of Defense, the Intelligence Community, and federal law enforcement...Contract workRemote workRelocation
- ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the operational excellence, resilience, and security of our core systems for a key Randstad client in the Washington D.C. area. This...
$149.4k - $202k
Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on...Remote work$115.5k - $164.8k
...that matters at a company where you matter. Your Impact As an engineer on the APX SRE CloudOps team, you will spend a significant... ...that replace what previously required human intervention with reliable, tested automation. You will also participate in on-call rotations...Full timeWork experience placementWork at officeRemote work- ...architecture of infrastructure and/or service according to terms for reliability and functionality. - Assists team members responding to... ...bottlenecks and deployments. - Gains basic knowledge of site reliability trends and shares relevant information with immediate...Full timeImmediate startFlexible hoursShift work
$84.24k - $142.48k
Overview Join us to work collaboratively with our talented team of dynamic and passionate engineers to deliver capabilities that enable our customers to make a difference. You'll deploy and operate ArcGIS Velocity and ArcGIS Workflow Manager SaaS solutions. You will also...WorldwideFlexible hours- ...an active U.S. Government Security Clearance at the TS/SCI level with required polygraph. We are seeking a Geospatial Platform Engineer to support geospatial, imagery, AI/ML, and data-driven application development, deployment, and operations. This role will focus on...Full timeRemote work
$116.9k - $243.1k
...Accenture Federal Services is seeking an experienced Release Train Engineer to join our team and support our client in the Northern VA area. The role requires dedicated and specialized support to accelerate Research, Development, Test, and Evaluation (RDT&E) efforts,...Live inLocal area$126k - $248k
...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and...Local areaRemote workWorldwideFlexible hours- ...position. Position Summary: ISI is looking for a Project Engineer Level 2 to provide Owner's Representative construction management... ...environments. Responsibilities: · Assist the Government in site evaluations, field surveys, and site visits to assess...Permanent employmentFor contractorsWork experience placementWork at officeMonday to Friday
- ...Senior PostgreSQL Database Reliability Engineer Responsibilities: Production PostgreSQL infrastructure across the full lifecycle: architecture, deployment, replication, monitoring, performance tuning, backup/recovery, and capacity planning Streaming replication...
- ...Job Title: Software Engineer - Senior Level (IE Platform Infrastructure) Location: Arlington, VA Clearance: Top Secret (TS) (Active... ...enterprise platform infrastructure. This role focuses on designing reliable integration environments, managing secure data exchange...Full timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site safety McLean, VA
- historic site McLean, VA
- IT site lead McLean, VA
- site leader McLean, VA
- website content developer McLean, VA
- junior website developer McLean, VA
- construction site safety McLean, VA
- official site McLean, VA
- on-site clinical research associate (traveling/remote) McLean, VA
- site reliability engineer sre





