Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$100k - $160k

Cathexis

Team CATHEXIS elevates the government contracting experience through rapid response, deep skill, and thoughtful problem-solving and communication. Our core capabilities are our top-tier program and project management, data analytics, and audit services, the backbone of which is our integrated approach to operational excellence.

You worked hard to get to where you are. You strive to make every day better than the day before. So do we. Team CATHEXIS operates with an all-in mindset. We are working together to create a company that supports our shared values and individual goals. Our values are centered around leading with integrity, owning the outcome, growing together, and moving with purpose in everything we do for our employees, customers, partners, and communities. We believe success is best when we listen and lead with empathy; model high standards of ethics to provide a rewarding candidate experience; work hard, have fun, and appreciate the strengths we all bring to the team; and empower our employees to create innovative and trusted results.

We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and optimize clusters on Kubernetes. Together, we're accelerating our clients' digital transformation through the building and deployment of data-driven, scalable AI solutions. The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will be responsible for ensuring the reliability and scalability of our clients' Kubernetes clusters and Cloud Infrastructure.

Responsibilities

The responsibilities include, but are not limited to:

  • Monitor and Manage Kubernetes Clusters: Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on Kubernetes
  • Kubernetes Management: Deploy, monitor, and scale applications on Kubernetes clusters. Maintain Helm charts, manage services, and ensure resource allocation for optimal cluster performance
  • Containerization & Deployment: Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environments
  • Cloud Infrastructure Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.)
  • Monitoring & Incident Response: Set up monitoring solutions, define alerts, an manage the incident response process for any issues related to Jenkins or Kubernetes clusters
  • Automate Infrastructure Processes: Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalent
  • Collaborate Across Teams: Work closely with development, services, and operations teams to ensure a seamless integration between application development, deployment, and infrastructure
  • Security & Compliance: Ensure all systems follow best practices in terms of security and compliance with relevant regulations. This includes role-based access, encryption, and automated vulnerability scanning

Team CATHEXIS elevates the government contracting experience through rapid response, deep skill, and thoughtful problem-solving and communication. Our core capabilities are our top-tier program and project management, data analytics, and audit services, the backbone of which is our integrated approach to operational excellence.

You worked hard to get to where you are. You strive to make every day better than the day before. So do we. Team CATHEXIS operates with an all-in mindset. We are working together to create a company that supports our shared values and individual goals. Our values are centered around leading with integrity, owning the outcome, growing together, and moving with purpose in everything we do for our employees, customers, partners, and communities. We believe success is best when we listen and lead with empathy; model high standards of ethics to provide a rewarding candidate experience; work hard, have fun, and appreciate the strengths we all bring to the team; and empower our employees to create innovative and trusted results.

We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and optimize clusters on Kubernetes. Together, we're accelerating our clients' digital transformation through the building and deployment of data-driven, scalable AI solutions. The ideal candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices. You will be responsible for ensuring the reliability and scalability of our clients' Kubernetes clusters and Cloud Infrastructure.

Responsibilities

The responsibilities include, but are not limited to:

  • Monitor and Manage Kubernetes Clusters: Ensure the stability, health, and scalability of Kubernetes Clusters, deploying applications and services on Kubernetes
  • Kubernetes Management: Deploy, monitor, and scale applications on Kubernetes clusters. Maintain Helm charts, manage services, and ensure resource allocation for optimal cluster performance
  • Containerization & Deployment: Design and maintain Docker-based microservices architecture, ensuring consistent and reproducible deployments across staging, QA, and production environments
  • Cloud Infrastructure Management: Work with leading Cloud Platforms (AWS, Azure and/or GCP) to set up, configure, and manage infrastructure resources using Infrastructure as Code (Terraform, CloudFormation, etc.)
  • Monitoring & Incident Response: Set up monitoring solutions, define alerts, an manage the incident response process for any issues related to Jenkins or Kubernetes clusters
  • Automate Infrastructure Processes: Build automation tools for scaling, monitoring, and maintaining infrastructure using modern tools like Terraform, Ansible, Linux, or equivalent
  • Collaborate Across Teams: Work closely with development, services, and operations teams to ensure a seamless integration between application development, deployment, and infrastructure
  • Security & Compliance: Ensure all systems follow best practices in terms of security and compliance with relevant regulations. This includes role-based access, encryption, and automated vulnerability scanning

  • Active TOP SECRET clearance or higher is required
  • Bachelor's degree in Computer Science or related field
  • A minimum of two (2) years of experience working with on-premise and off-premise cloud environments
  • Experience with AWS and/or Azure
  • Hands-on experience with a range of open-source technologies, such as Linux, Docker, Kubernetes, K8s, Terraform, Helm, PostgreSQL, or similar technologies
  • Ability to program (structured and OOP) using one or more high-level languages, such as Python, Java, C/C++, Ruby, and JavaScript
  • Experience with distributed storage technologies such as NFS, HDFS, Ceph, and Amazon S3, as well as dynamic resource management frameworks (Apache Mesos, Kubernetes, Yarn)
  • Proactive approach to identifying problems, performance bottlenecks, and areas for improvement
  • Ability to lead and work independently in an Agile/Scrum environment
  • Real passion for developing team-oriented solutions to complex engineering problems
  • Thrive in an autonomous, empowering and exciting environment
  • Great verbal and written communication skills to collaborate multi-functionally and improve scalability
  • Interest in committing to a fun, friendly, expansive, and intellectually stimulating environment
Desired Skills
  • Hands-on experience deploying and operating applications using IaaS and PaaS on major cloud providers, such as Amazon AWS, Microsoft Azure, or Google Cloud Services
  • Experience with deep learning, natural language processing, computer vision, or reinforcement learning
  • Conveys highly technical concepts and information in written form to technical and non-technical audiences
  • The ability to work on multiple concurrent projects is essential. Strong self-motivation and the ability to work with minimal supervision
  • Must be a team-oriented individual, energetic, result & delivery oriented, with a keen interest on quality and the ability to meet deadlines

CATHEXIS offers competitive compensation packages to all eligible employees. Our goal is to provide a compensation package that reflects the value you bring to our team, is competitive with national average market rates, and promotes your financial security and personal well-being. The annual salary range for this role is $100,000 - $160,000. Please note that the salary information provided is a general guideline. CATHEXIS considers various factors in its final offer, including location, qualifications, experience, and skills.

  • Performance Bonuses
  • Medical Insurance
  • Dental Insurance
  • Vision Insurance
  • 401(k) Plan (Traditional and ROTH)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off
  • 11 Federal Holidays
  • Parental Leave
  • Commute Benefits
  • Short Term & Long Term Disability
  • Training & Development
  • Wellness Program
  • Community Outreach Initiatives

CATHEXIS is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability EEO IS THE LAW. If you are an individual with a disability and would like to request a reasonable accommodation as part of the employment selection process, please contact the Recruiting View email address on click.appcast.io

Vacancy posted 22 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in McLean, VA vacancy
  •  ...Site Reliability Engineer Description: We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer (SRE) will manage, monitor, and optimize clusters on Kubernetes. Together, we're accelerating... 
    Suggested

    STEM Solutions

    McLean, VA
    21 hours ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Suggested
    Remote work

    Govcio LLC

    Arlington, VA
    4 days ago
  • $120k - $150k

     ...operates along the way. In This Role, You Will: Own reliability across a complex, global product portfolio. You'll be...  ...Daily use of Claude Code, Cursor, or AI coding assistants as an engineering accelerant - not occasional dabbling Experience... 
    Suggested
    Work at office
    Worldwide

    Cvent

    McLean, VA
    2 days ago
  • $106.3k - $221.1k

     ...Senior Site Reliability Engineer At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life better for people. Our 13,000+ people are united in a shared purpose to pursue the limitless potential... 
    Suggested

    Accenture Federal Services

    Arlington, VA
    4 days ago
  •  ..., to act first. Exiger is FedRAMP authorized and a 2x Leader in Gartner Magic Quadrant for Supplier Risk Management. Site Reliability Engineer Location: U.S. (Hybrid) This role requires U.S. citizenship and eligibility for a U.S. security clearance. Role... 
    Suggested
    Work at office
    Work from home
    Flexible hours

    Exiger

    McLean, VA
    22 hours ago
  • $128.5k - $190k

     ...exceptional people to create extraordinary experiences together. Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure and applications that power a highly reliable global SaaS... 
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    1 day ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track Key Performance Indicators (KPIs) and Service Level Objectives (SLOs), identify and resolve performance bottlenecks, and perform... 
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    2 days ago
  •  ...Catalyst, and In-Q-Tel. Mission | On Site | Full Time | Active TS/SCI with Full...  ...government customer site, ensuring the reliability and performance of Twenty's mission-critical...  ...technical ownership and customer-facing engineering: you'll define how we measure... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Technologies

    Arlington, VA
    2 days ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    GovCIO

    Arlington, VA
    4 days ago
  • $150k - $180k

     ...what’s possible in remote sensing, you belong here at Umbra. About the Job We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's systems.... 
    Permanent employment
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    Umbra

    Arlington, VA
    22 hours ago
  • $106.3k - $221.1k

     ...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key... 
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    22 hours ago
  • $90k - $130k

     ...critical, world-changing federal challenges. Credence has an immediate opening for a Site Reliability SME who has hands-on experience working as a Cloud Operations Engineer with experience in IT operations to join our expanding Cloud Managed Services Provider team... 
    Temporary work
    Work experience placement
    Immediate start
    Worldwide

    Credence

    McLean, VA
    12 days ago
  • $158.5k - $230k

     ...together. Bring your whole self. The Role and Team We are growing our GovCloud team and looking for a Staff Site Reliability Engineer to help scale how we operate Medallia's US public-sector cloud platform. You will support federal agencies and other... 
    Permanent employment
    Temporary work
    Work experience placement
    Work at office
    Local area
    Remote work
    3 days per week

    Medallia

    McLean, VA
    4 days ago
  • $145k - $160k

     ...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep... 
    Temporary work
    Remote work
    Flexible hours

    EPAM Systems Inc

    McLean, VA
    22 hours ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    a month ago
  •  ...Job Description Job Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient... 
    Local area

    Tiger Analytics Inc.

    Washington DC
    a month ago
  •  ...Site Reliability Engineer (SRE) Reston, VA Site Reliability Engineer (SRE) Position: Site Reliability Engineer (SRE) Work Authorization: All Work Authorizations Location: Reston, VA Contract: 24 months Description: Site Reliability Engineer (SRE) roles... 
    Contract work

    Knack Solutions

    Reston, VA
    2 days ago
  • $81.1k - $187k

     .... You'll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers...  .... Responsibilities Escalation points for junior Site Reliability Engineers during complex or high-impact incidents. Manage... 
    Temporary work
    Work experience placement
    Monday to Friday
    Flexible hours
    Shift work
    Night shift

    Oracle

    Reston, VA
    3 days ago
  •  ...solutions using a tailored Agile methodology. We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the... 
    Remote work

    Elevate Government Solutions

    Washington DC
    1 day ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Work experience placement

    Samprasoft

    Washington DC
    22 hours ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Reston, VA
    22 hours ago
  •  ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington D.C. area, focusing on optimizing the availability, performance, and scalability of critical production services. The ideal... 

    Software Technology Inc

    Washington DC
    4 days ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer Location- Wilmington De, Washington DC, Dallas, TX (Onsite Position) Full time position Minimum Qualifications Bachelor’s degree in computer science, Engineering, or a related technical field. Minimum of 5 years of experience... 
    Full time

    Yochana

    Washington DC
    1 day ago
  •  ...Site Reliability Engineer Location: Occasional onsite visits to Reston VA (Zip code 20190). Duration-1 year plus Interview process: The final interview is a mandatory, face-to-face interview in Reston VA. Zip code: 20190 Strong... 
    Long term contract
    Temporary work
    H1b
    Immediate start
    Relocation

    3B Staffing LLC

    Reston, VA
    4 days ago
  • $135k - $154k

     ...where you matter.Your ImpactAs a contributor in the APX platform engineering organization on the CloudNet team, you are passionate about...  .... You are also obsessed about achieving the high quality and reliability our customers demand. You will work closely with sovereign... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    4 days ago
  • $121.4k - $218.6k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    3 days ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    3 days ago
  • $160k - $180k

     ...Site Reliability Engineer Location: Hybrid – Washington DC/Virginia/Maryland metro with the ability to travel to Patuxent River, MD, as needed (up to 20% of the time). Compensation: $160,000 - 180,000 per year, depending on experience and qualifications. Employment... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Fortress Information Security

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise Cloud transformation... 
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    22 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!