Reliability Engineer
GAP Solutions, Inc. (GAPSI)
Position Objective: In this role, the Reliability Engineer will ensure the reliability, scalability, and operational health of EDAV’s Azure cloud environment. Terraform is central to this position: the engineer will independently design, build, review, and troubleshoot Infrastructure-as-Code for mission critical environments. The role involves close collaboration with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve Kubernetes operations, and resolve issues impacting the availability of EDAV data and analytics services.
Duties and Responsibilities:
Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.
Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).
Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.
Automate cloud operations to reduce manual work and improve consistency.
Collaborate with development and operations teams to enhance deployment and incident response practices.
Implement and refine monitoring, alerting, dashboards, and operational reporting.
Identify reliability risks and recommend improvements to cloud architecture and processes.
Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.
Qualifications
Basic Qualifications:
4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles
2+ years of hands‑on Microsoft Azure experience
2+ years of hands‑on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructure
Experience administering/supporting Kubernetes (preferably AKS)
Experience supporting cloud‑hosted systems and automating cloud operations
Knowledge of cloud networking concepts and troubleshooting
Possession of strong analytical and problem‑solving abilities
Ability to work on-site in Atlanta, GA or the Washington D.C. metro-area
Ability to obtain/maintain Public Trust/Suitability clearance
Bachelor’s degree or equivalent experience
Must be U.S. Citizen or Lawful Permanent Resident (Green Card Holder)
Preferred Qualifications:
Experience with certificate lifecycle management (maintaining, renewing, rotating certs)
Experience creating operational dashboards or reports (Power BI preferred)
Experience with observability platforms (Grafana, Prometheus, Elastic, Splunk)
Experience integrating Azure resources with Active Directory
Experience using agentic coding tools (Claude Code, Codex, GitHub Copilot) for applications or infrastructure automation
Azure, Kubernetes, or Terraform certifications
* This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required by this position.
To perform this job successfully, an individual must be able to perform each essential duty satisfactorily. The requirements listed above are representative of the knowledge, skill, and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
GAP Solutions provides reasonable accommodations to qualified individuals with disabilities. If you need an accommodation to apply for a job, email us at View email address on click.appcast.io . You will need to reference the requisition number of the position in which you are interested. Your message will be routed to the appropriate recruiter who will assist you. Please note, this email address is only to be used for those individuals who need an accommodation to apply for a job. Emails for any other reason or those that do not include a requisition number will not be returned.
Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, protected veteran status or other characteristics protected by law.
$68 - $75 per hour
DescriptionWe are seeking a Reliability Engineer for our customer at the CDC. In this role, the engineer will ensure the reliability, scalability, and operational health of EDAV’s Azure cloud environment. Terraform is central to this position: the engineer will independently...SuggestedContract workTemporary work- Job DescriptionJohnson Service Group (JSG) is seeking a qualified Reliability Engineer in Atlanta, GA. This is an opportunity to work for a growing company. Role Summary Nexus Circular is seeking a highly analytical and experienced Contract Reliability Engineer to lead...SuggestedWeekly payContract workTemporary work
$95k
...Exceptional Performance ! Chenega Services & Federal Solutions, LLC , a Chenega Professional Services company, is looking for a Reliability Engineer. In this role, the Reliability Engineer will ensure the reliability, scalability, and operational health of EDAVs Azure...SuggestedContract work- ...LinkedIn Tag How you'll help us Keep Climbing (overview & key responsibilities) As a member of the Reliability Programs team, the Reliability Engineer supports their assigned Fleet Engineering team and reports directly to the manager of Reliability Programs. In...SuggestedFull timeTemporary workWork at officeRemote work
$126k - $167k
...Reliability Engineer Atlanta, Georgia, United States; Quonset, Rhode Island, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology...SuggestedFull timeWork experience placementImmediate start$98.18k - $115.5k
...alertgovernance,andmonitoringbestpractices.ServeasatrustedadvisortoProduct,Engineering,SRE,Infrastructure,andOperationsleaders,... ...Skills/ExperienceExpertise in Observability Engineering, Site Reliability Engineering (SRE), or Reliability Engineering.Strong knowledge...Full timeWork experience placementLocal area3 days per week- ...Job Description Job Description TRC is hiring a Reliability Engineer for our client a nationally recognized restaurant organization where you'll play a critical role in improving equipment reliability across nearly 3,000 locations. In this highly visible position,...Permanent employmentWork at office
- ...Description Job Description Are you passionate about ensuring the reliability and performance of advanced semiconductor technologies? At... ...systems. Falcomm is seeking an RFIC Reliability Engineer to lead reliability analysis and qualification activities for...Permanent employmentFull time
- ...Role Summary Focus on improving equipment reliability and reducing maintenance issues across approximately 3,000 restaurant locations... ...identify root causes, and drive corrective actions. Lead engineering improvements that reduce downtime, repair costs, and...Permanent employmentFull timeWork at office
$113k - $171.6k
...are growing rapidly and hiring top talent with leading AI skills across engineering, sales, product, marketing, and beyond as we build the leading digital operations platform.As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office,you'll...Work at officeLocal areaFlexible hours- ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
- OverviewJob PurposeAt Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect... ...results-oriented people to join our team.We are seeking a Site Reliability Engineer to bring 3+ years of hands-on experience to our SRE team...
$104.9k - $174.7k
...Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role...Full timeWork at officeLocal areaRemote workWork from home- #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...
$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...Full timeTemporary workWork experience placementFlexible hours$141.3k - $237.4k
...AT&T, you won’t just imagine the future, you’ll build it.We are seeking a highly skilled and hands-on Lead Software Engineer to join Software Reliability Engineering (SRE) Onboarding and automation team. This role will drive innovation through automation, enhancement of...Full timeTemporary workWork at officeLocal areaRelocation$130k - $150k
...to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will have...Full timeRemote work- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...
$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential and is located in Atlanta, GA (Onsite). Please review the job description below and contact me ASAP if you are interested...Contract workLocal areaImmediate start- ...Site Reliability Engineer (SRE) When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage...Work at office
- ...and we're looking for new team members who want to be a part of this journey! We're looking for a proactive, hands-on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in fast-moving startup environments. You're someone who enjoys...Work experience placementFlexible hours
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...We are currently looking for a Senior Software Engineer to be a part of the Site Reliability Engineering (SRE) team in Atlanta, GA . The SRE team is an innovative team devoted to providing a Docker-based Platform as a Service and assisting a growing number of teams...Contract workWork at officeLocal area
- ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and...
$120k - $175k
...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues...Full timeRemote workWork visaFlexible hours- ...Site Reliability Engineer Seek Now is transforming property inspections through technology, data, and human expertise. We deliver faster, smarter, more reliable insights to insurance carriers and single-family rental markets, and we're just getting started. If you...Flexible hours
$55 - $60 per hour
...Description Our client is looking for an SRE that will Lead the reliability, scalability, security, and operational excellence of... ...solutions to improve operational efficiency. • Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance...- ...Join to apply for the Site Reliability Engineer role at Motion Recruitment Join to apply for the Site Reliability Engineer role at Motion Recruitment Get AI-powered advice on this job and more exclusive features. Every year, nearly 200 million travelers...Contract workWorldwide
- ...Site Reliability Engineer At Acuity, you will join an Agile team focused on building and supporting advanced platforms and applications that drive our business forward. We are seeking a Site Reliability Engineer (SRE) to help define and raise the reliability bar for...
- ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively...Work experience placement
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Reliability Engineer. Be the first to apply!
- reliability maintenance engineering technician Atlanta, GA
- reliability engineer Atlanta, GA
- principal reliability engineer
- senior reliability engineer
- hardware reliability engineer
- reliability engineering manager
- reliability maintenance engineering technician
- maintenance & reliability engineer
- reliability engineer
- sr reliability engineer


