Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Exiger

Who We Are:

Exiger transforms supply chains into a strategic advantage, advancing our mission to make the world a safer and more transparent place to succeed. OurAI platform, 1Exiger, delivers instant visibility into complex supplier ecosystems, leveraging proprietary data and advanced AI to surface risk, automate compliance, and unlockefficiencies and cost savings to strengthen long-term resilience. Trusted by 550+ global customers, including Fortune 500 companies and U.S. government agencies, Exiger is a recognized, award-winning leader in supply chain AI and a FedRAMP authorized provider to the federal government.

Site Reliability Engineer

Location: U.S. (Hybrid)

This role requires U.S. citizenship and eligibility for a U.S. security clearance.

Role Summary:

Exiger is transforming how governments and global enterprises manage supply chain, defense, and geopolitical risk. Our AI-powered platform equips the world's most important institutions with the intelligence they need to protect critical infrastructure, secure national interests, and make data-driven operational decisions.

From identifying counterfeit parts in defense supply chains to anticipating geopolitical risk exposure, Exiger enables mission owners to act with clarity and confidence in complex, high-stakes environments.

This is our first dedicated Site Reliability Engineering hire and a founding role. You will help stand up the SRE function at Exiger: setting the standards, tooling, and practices that keep 1Exiger reliable for our 550+ customers, including Fortune 500 companies and U.S. government agencies. You will own reliability across the full service lifecycle, from design and capacity planning through deployment, monitoring, and incident response, and build the automation that lets the platform scale without scaling headcount. Because you are first, we need someone who has practiced SRE before and can bring the playbook, not learn it on the job.

You will use your expertise in coding, algorithms, complexity analysis, and large-scale distributed system design to solve the reliability challenges that are unique to operating a mission-critical AI platform in regulated and government environments.

SRE's culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

What You'll Do:

  • Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards that other engineering teams adopt.
  • Build and own observability: instrument services for availability, latency, and system health, and turn that signal into actionable insight.
  • Drive decisions with data: form hypotheses, measure the impact of every change, and let metrics rather than intuition set reliability priorities.
  • Own the reliability of production services from design consulting and launch reviews through steady-state operation.
  • Eliminate repetitive manual operations through automation and infrastructure as code, replacing them with reliable, self-service tooling.
  • Plan for scale: capacity planning, performance analysis, and driving changes that improve both reliability and delivery velocity.
  • Improve resilience through chaos engineering and fault-injection testing, running game days that prove the platform degrades gracefully and recovers from failure.
  • Lead sustainable, blameless incident response and postmortems, and stand up and participate in an on-call rotation.
  • Leverage AI-assisted development tooling (such as Codex and Claude) to accelerate automation, tooling, and investigation work, and help the team adopt these tools effectively.


What You Need:

  • Bachelor's or Master's degree in Computer Science, a related field, or equivalent practical experience.
  • 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role. As our first SRE hire, you must have practiced SRE before and be ready to establish the function.
  • 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems.
  • Strong grounding in Unix/Linux internals (filesystems, processes, system calls) and networking fundamentals (TCP/IP, DNS, routing, load balancing).
  • Hands-on experience establishing core SRE practices from the ground up: SLIs, SLOs, and error budgets, monitoring and observability, capacity planning, and automation that removes repetitive manual work.
  • A rigorous, empirical mindset: you form hypotheses, measure outcomes, and make metrics-driven decisions rather than relying on intuition or anecdote.
  • Experience with chaos engineering or fault-injection testing (for example game days, Chaos Monkey, Gremlin, or LitmusChaos) to validate system resilience.
  • Proven incident management experience: on-call ownership, leading response under pressure, and driving blameless postmortems to root cause.
  • Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems.
  • Familiarity with cloud platforms (AWS) and secure system integration.
  • Comfort integrating AI coding assistants (such as Claude and Codex) into your daily engineering workflow.
  • Ability to translate ambiguous mission problems into structured technical solutions.
  • Ability to operate independently in dynamic, high-stakes environments.
  • Willingness to travel as needed to support customer engagements.


Nice to Have:

  • 4 years of experience programming in Go or C (Java also welcome), with the ability to debug, optimize, and automate rather than just script.
  • Experience supporting ML or data platforms in production.
  • Familiarity with data warehouses such as Snowflake, Redshift and/or Apache Iceberg.
  • Experience operating in FedRAMP or other regulated or government environments.


Why You'll Love Working at Exiger:

  • High-performance culture rooted in accountability, collaboration, and a shared commitment to excellence.
  • Discretionary Time Off for all employees, with no maximum limits on time off
  • Industry leading health, vision, and dental benefits
  • Competitive compensation package
  • 16 weeks of fully paid parental leave
  • Flexible, hybrid approach to working from home and in the office where applicable
  • Focus on wellness and employee health through stipends and dedicated wellness programming
  • Purposeful career development programs with reimbursement provided for educational certifications


#LI-hybrid

Exiger is named a Leader in the GartnerMagic Quadrant for Supplier Risk Management, twice selected as one of Fast Company's 'Brands That Matter,' and recipient of the Third Party Risk Association's Innovator Award, Exiger's technology has been recognized by leading analyst evaluations and 50+ awards. Learn more at Exiger.com and follow Exiger on LinkedIn .

At Exiger, our values define how we work and why we lead. We are mission-inspired, imagination-driven, trust-anchored, and compassion-focused-committed to building technology that makes the world safer, more transparent, and more resilient.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law.

Exiger's hybrid work policy is periodically reviewed and adjusted to align with evolving business needs.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Richmond, VA vacancy
  • Job Title Top skills: ~ Kubernetes ~ CKA certification is must ~ New relic and Splunk - document the issues and provide resolution, troubleshooting and creating new alerts ~3+ yrs of experience as SRE with Kubernetes
    Suggested

    Samprasoft

    Richmond, VA
    1 day ago
  • $121.4k - $218.6k

     ...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with...  ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust... 
    Suggested
    Work experience placement
    Work at office

    Akamai

    Richmond, VA
    4 days ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Suggested
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Richmond, VA
    2 days ago
  •  ...About the job Site Reliability Engineer ***W2 only*** Position: Site Reliability Engineer (SRE) Work Authorization: All Work Authorizations Location: Richmond, VA Contract: 24 months As one of the Site Reliability Engineers, youll be able to work... 
    Suggested
    Contract work
    Immediate start

    Knack Solutions

    Richmond, VA
    3 days ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Suggested
    Work experience placement
    Work at office

    Akamai

    Richmond, VA
    4 days ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Richmond, VA
    4 days ago
  • $100k - $130k

     ...driven and highly collaborative team that is passionate about creating transformative change in healthcare. We are seeking a Site Reliability Engineer to play a key role in designing, optimizing, and securing the underlying cloud infrastructure that powers our organization... 
    Flexible hours

    Datavant

    Richmond, VA
    5 days ago
  • $74.1k - $148.3k

     ...systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by defining, designing, deploying, and solving key Oracle Cloud services,... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Richmond, VA
    1 day ago
  •  ...Lead Site Reliability Engineer McLean, VA or Richmond, VA - onsite position - will consider relocations( But try to find Local) Must have senior level experience with the following: Strong reliability and operational mindset (not frontend development... 
    Local area
    Relocation

    3B Staffing LLC

    Richmond, VA
    7 hours ago
  • $84.9k - $209.5k

     ...Job Description As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve... 
    Temporary work
    Flexible hours

    Oracle

    Richmond, VA
    1 day ago
  • $96.7k - $145k

     ...world and each other. \n POSITION TITLE \n Senior Platform Engineer – AWS Foundations \n POSITION LOCATION  \n Richmond, VA or Lynchburg...  ...the foundational AWS services that power secure, scalable, and reliable application delivery within the AWS ecosystem. \n This role... 
    Work at office
    Local area
    Remote work

    Genworth

    Richmond, VA
    4 days ago
  • $75.4k - $126.5k

     ...thinking.Woolpert is an award-winning, global leader in architecture, engineering, and geospatial services. We blend design excellence with...  ...for career growth.Position OverviewWoolpert is hiring a Site Civil Engineer to join our dynamic Land Development team. This... 
    Casual work
    Local area
    Remote work
    Flexible hours
    Night shift

    Woolpert

    Richmond, VA
    1 day ago
  •  ...JOB TITLE: Release Train Engineer JOB LOCATION: Remote but must be open to travel to Alpharetta, GA when needed WAGE RANGE*: $45 to $48hr JOB NUMBER: 26-00997 REQUIRED EXPERIENCE: Minimum of five years Program Management experience and/or five years of experience as a... 
    Hourly pay
    Temporary work
    Local area
    Remote work
    Flexible hours

    Computer Merchant, Ltd., The

    Richmond, VA
    4 days ago
  •  ...be comfortable with autonomously leading in a matrixed environment. Required Skills: ~5 Years - SAFe Agile Release Train Engineer Hands-on Experience. ~3 Years - SAFe Agile Release Train Engineer Experience managing Power Platform teams. ~5 Years -... 

    3B Staffing LLC

    Richmond, VA
    7 hours ago
  •  ...preferred. About Delphi-US Delphi-US is a national recruiting firm based in Newport, Rhode Island. We specialize in IT, Engineering and Professional Staffing services for organizations across the United States of America.? Our mission is simple: We are... 
    Contract work
    For contractors
    Work experience placement

    Delphi-US

    Richmond, VA
    7 hours ago
  •  ...Reliability Engineer Responsibilities Performs in-depth root cause analysis (RCA) of equipment failures. To this end, bring together a committee of peers and guide the process. Manage and prioritize projects and RCAs, ensuring effective planning and execution... 

    Hatch Global Search

    Richmond, VA
    3 days ago
  • $103.71k - $138.28k

     ...demonstrated knowledge and experience in system architecture and engineering disciplines. Specific technical knowledge of enterprise level...  ...Amazon Web Services. -Supports due diligence activities including site surveys, design, design review, bill of materials creation,... 
    Temporary work
    Remote work

    Lumen Inc

    Richmond, VA
    3 days ago
  • $90k - $125k

     ...and government agencies, a list that spans across the country.Job DescriptionThe Opportunity: CapTech is seeking a SaaS platform engineer to play a hands-on role in the engineering, administration, and evolution of our enterprise application ecosystem. This is not a traditional... 
    Visa sponsorship
    Work visa

    CapTech

    Richmond, VA
    3 days ago
  • $157.03k - $240.14k

    THE IMPACT YOU WILL HAVE: As a QTS Senior Solutions Engineer for our Enterprise Strategic Accounts team, you will serve as the primary technical resource supporting both sales and operations. You will identify and define customer requirements throughout the evaluation... 
    Full time
    Work at office

    Quality Technology Services

    Richmond, VA
    2 days ago
  •  ...MITRE Systems Engineering Role Why choose between doing meaningful work and having a fulfilling life? At MITRE, you can have both. That...  ...executives or senior customers. Ability to travel to sponsor sites at Pentagon and Mark Center bi-weekly. Candidate must... 
    Temporary work
    Work experience placement
    Internship
    Work at office
    Local area

    MITRE

    Richmond, VA
    2 days ago
  •  ...learning from the world and each other.POSITION TITLESenior Software Engineer (Full Stack)This role is not eligible for employment visa...  ..., and runtime operations by improving observability, reliability, and performance (monitoring, alerting, incident response, and... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work

    Genworth Financial

    Richmond, VA
    4 days ago
  • $80 - $95 per hour

     ...complexity applications ensuring that the code follows latest coding practices and industry standards. Work closely with other software engineers and team members to understand the system end-to-end and perform system analysis. Analyze, design, configure and develop the... 
    Contract work
    Temporary work
    Remote work

    TEKsystems

    Richmond, VA
    3 days ago
  • $118.4k - $219.8k

    Are you a hands-on AI engineer who wants to build the next generation of legal AI solutions from the ground up?Thomson Reuters is hiring a Senior Software Engineer - AI on our CoCounsel Forward Deployed Engineering team. In this role, you will help design, build, and deploy... 
    Full time
    Contract work
    Local area
    Flexible hours

    Thomson Reuters

    Richmond, VA
    2 days ago
  •  ...Electrical Reliability Engineer We are seeking an Electrical Reliability Engineer to join our Engineering Team to provide leadership for electrical...  ...to advance R&D programs and commercial performance. Be a site technical leader and subject matter expert for electrical... 

    MRINetwork

    Richmond, VA
    7 hours ago
  • $174k - $236k

     ...experience and joining our team?About the roleWe're hiring a hands-on engineering leader to own both the people and the delivery of our platform...  ...helps engineers growBuild a team culture that is organized, reliable, and focused on impactMentor engineers through code reviews,... 
    Work at office
    Remote work
    Flexible hours

    Koalafi

    Richmond, VA
    5 days ago
  • $33.55 - $39.47 per hour

    Job Title Lead Building Engineer Job Description Summary Responsible to ensure the proper...  ...developed to assure maximum life and reliability of all mechanical/ electrical/plumbing systems...  ...does not have a Chief Engineer on-site at the building and is sometimes the solo... 
    Minimum wage
    For contractors
    Apprenticeship
    Work at office
    Immediate start
    Flexible hours
    Shift work
    Weekend work
    Day shift

    Cushman & Wakefield

    Richmond, VA
    4 days ago
  • $97.3k - $146k

     ...learning from the world and each other.POSITION TITLEPlatform Engineer, Azure / DevOpsPOSITION LOCATIONThis position is hybrid and located...  ...and optimize resources to ensure high availability and reliability.• Monitor and analyze query performance and optimize data storage... 
    Full time
    Local area
    Flexible hours

    Genworth Financial

    Richmond, VA
    4 days ago
  • Syms Strategic Group (SSG)  is seeking a talented Software Developer Location: Remote Department: Veterans Affairs Type:  Full Time Min. Experience:  Experienced Security Clearance Level:  Public Trust (MBI)  Military Veterans are highly encouraged...
    Full time
    Remote work

    Ssg

    Richmond, VA
    7 hours ago
  • $63.72k - $110.65k

     ...Description Job Description As a global leader in consulting, engineering, and commissioning services, we specialize in highly technical...  ...firm is active within. The position might require travel to sites throughout the US and provide the opportunity to interface... 
    For contractors
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours

    Syska Hennessy Group

    Richmond, VA
    8 days ago
  •  ...Please review the following job description:The Senior Security Engineer on the Proxy Team is responsible for evolving and sustaining the...  ...details on Truist’s generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be... 
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Work visa
    Shift work
    Day shift

    Truist

    Richmond, VA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!