DevOps / Site Reliability Engineer
AgileEngine
Job Description
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards. WHY JOIN US If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment. WHAT YOU WILL DO - Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP). - Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations. - Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment. - Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads. - Serve as Incident Commander on major and critical incidents - running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure. - Own the post-incident loop - track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups. - Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle. - Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution. MUST HAVES - You must be authorized to work for ANY employer in the US (e.g., Green card holders, TN visa holders, GC EAD, H4 EAD, U4U with EAD), as we are unable to sponsor or take over employment visa sponsorship at this time; - 5+ years of experience . - In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles . - Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting . - Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment. - Proven track record of remediation follow-up - coordinating with teams and holding owners accountable until issues are fully closed. - Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences. - Direct experience authoring divisional/group incident-management playbooks and escalation procedures. - Fully autonomous. - Drives the architecture of complex automated runbooks and mentors Middle-level SREs. - Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz . - Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2) . - Upper-intermediate English level. NICE TO HAVES - PagerDuty - hands-on experience with on-call scheduling, alert routing, and incident orchestration. - ServiceNow - familiarity with incident, problem, and change management workflows and reporting. PERKS AND BENEFITS - Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget - Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews - Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm - Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands - Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized - Well-being & support : access local well-being programs and people-focused support tailored to your location
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards. WHY JOIN US If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment. WHAT YOU WILL DO - Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP). - Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations. - Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment. - Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads. - Serve as Incident Commander on major and critical incidents - running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure. - Own the post-incident loop - track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups. - Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle. - Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution. MUST HAVES - You must be authorized to work for ANY employer in the US (e.g., Green card holders, TN visa holders, GC EAD, H4 EAD, U4U with EAD), as we are unable to sponsor or take over employment visa sponsorship at this time; - 5+ years of experience . - In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles . - Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting . - Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment. - Proven track record of remediation follow-up - coordinating with teams and holding owners accountable until issues are fully closed. - Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences. - Direct experience authoring divisional/group incident-management playbooks and escalation procedures. - Fully autonomous. - Drives the architecture of complex automated runbooks and mentors Middle-level SREs. - Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz . - Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2) . - Upper-intermediate English level. NICE TO HAVES - PagerDuty - hands-on experience with on-call scheduling, alert routing, and incident orchestration. - ServiceNow - familiarity with incident, problem, and change management workflows and reporting. PERKS AND BENEFITS - Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget - Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews - Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm - Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands - Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized - Well-being & support : access local well-being programs and people-focused support tailored to your location
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the DevOps / Site Reliability Engineer in Atlanta, GA vacancy
$104.9k - $174.7k
...link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability... ...CI/CD pipelines and operational workflows using Azure DevOps and GitHubWork directly with application teams to improve reliability...DevopsFull timeWork at officeLocal areaRemote workWork from home- ...Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation,... ...Engineering, Cloud Engineering, Platform Engineering, DevOps, or related technical disciplines.Demonstrated success designing...DevopsFull timeWorldwideFlexible hours
$152.13k - $162.13k
...about what’s next. Join us.General Summary:Unum Group seeks Site Reliability Engineers in Atlanta, GA.Applicants who are interested in this... ...engineering, platform, infrastructure, or operations teams, within a DevOps or reliability-focused environment; working with version...DevopsFull timeTemporary workWork at officeRemote work$138.1k - $198.2k
...technology that simply works. The SRE Engineering Enablement Team supports our CI Platforms... ...engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter... ...Familiarity with DORA and SPACE published DevOps measurement metrics. Deep interest in...DevopsPermanent employmentFull timeTemporary workWork experience placementLocal areaRemote workFlexible hours$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term... ...to achieve the automation AWS Pipeline & Infrastructure, DevOps (SRE) activities, Monitoring & Alerting Our client is...DevopsContract workLocal areaImmediate start- ...Site Reliability Engineering LeadThe Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise... ..., microservices, container orchestration, and DevOps, strong familiarity with Agile frameworks, continuous integration...Devops
- ...Lead Engineer, Site Reliability Engineering Team As a lead engineer with Retail, Site Reliability Engineering team, you will be at the forefront... ...experience 2+ years support a production system on a DevOps team 2+ years of experience running and building systems in...Devops
$167.7k - $245.2k
...cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly... ...Production Engineering, SRE, Site Reliability Engineering, DevOps, CCNP, CCIE, JNCP, JNCIE Why Cisco? At Cisco, we’re revolutionizing...DevopsPermanent employmentFull timeTemporary workLocal areaFlexible hours$120k - $175k
...need at least 5 years of experience as a reliability-focused engineer in a fast-moving, rapidly expanding... ...Python Ruby Security Terraform DevOps Windows More: We are... ...applicants anywhere in the U.S. This Senior Site Reliability Engineer role offers a...DevopsFull timeRemote workVisa sponsorshipFlexible hours- ...monitoring and alerting metrics so the support engineers can proactively and timely validate,... ...Experience • 3+ years of related DevOps, SysOps engineering experience with... ...application components. • 1+ Years in Site Reliability Engineering organization preferred •...DevopsWork experience placement
- ...of America)Please review the following job description:The Site Reliability Engineering Lead role focuses on enhancing the reliability and... ...architectures, microservices, container orchestration, and DevOps.4. Strong familiarity with Agile frameworks, continuous integration...DevopsPermanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross... ...solutions that build integrations between Dynatrace, Azure DevOps and Jira. Solid knowledge in focused areas of OneStream Software...DevopsFull timeTemporary workWork experience placementRemote work- ...System Reliability Engineer (SRE)At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give... ...designs and builds solutions using Power Platform, Azure DevOps Pipelines, Microsoft Graph APIs, and AI-powered automation to...DevopsFull timeTemporary workPart timeWork experience placementFlexible hours
- ...Senior Site Reliability Engineer Atlanta, Georgia Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered... ...industry experience ~7+ years of experience as a DevOps, SRE or Systems Engineer ~ Advanced proficiency with at least...DevopsPermanent employmentFull timeWork at officeRemote workWork from homeWork visa
- ...Manager @ STAFFWORXS | US IT Recruitment Job Opening: AWS Site Reliability Engineer (SRE) We’re hiring a Site Reliability Engineer (SRE) to join... ...& CI/CD: AWS CDK, GitLab CI Scripting & Automation: Python DevOps Practices & Tooling What We’re Looking For: Solid experience...DevopsContract work
$71.6k - $119.4k
...automation, troubleshoot issues, and work closely with senior engineers to learn and apply best practices. You'll gain exposure to a wide... ...audiences. Requirements ~1–3 years of experience in DevOps, SRE, cloud engineering, or related IT roles (internships and projects...DevopsTemporary workInternshipLocal area- ...place. Job Details: Job Title: SRE Engineer Location: Atlanta GA (Hybrid... ...Minimum Experience ~3+ years of related DevOps, SysOps engineering experience with focus... ...application components. ~1+ Years in Site Reliability Engineering organization preferred. ~...DevopsContract workWork experience placementWork at officeRemote work
- ...initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include... ..., and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps practices, raising the bar on what's possible...DevopsWork at officeLocal areaRemote workWork from homeWorldwide
- ...thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability... ...projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role. We care far more about...DevopsFull timeLive inWork at office
$35 - $45 per hour
DescriptionKforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot of Java/API'...Remote work$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...Full timeTemporary workWork experience placementFlexible hours- ...supports the Subscription Product Engineering organization, including in-... ...automation, monitoring, and reliability-focused practices across... ...way!Job Responsibilities:Apply DevOps automation tools to manage CI... ...practices (Required)Knowledge of Site Reliability Engineering...DevopsFull timeTemporary workPart timeWork experience placementLocal areaFlexible hours
$141.3k - $237.4k
...are seeking a highly skilled and hands-on Lead Software Engineer to join Software Reliability Engineering (SRE) Onboarding and automation team. This role... ...systems.Hands-on experience with CI/CD pipelines, DevOps practices, and Cloud architectureFamiliarity with observability...DevopsFull timeTemporary workWork at officeLocal areaRelocation- ...Release Train Engineer (RTE) At Xebia we are always looking for talented people to deliver value to our clients and their customers... ...experience Prior airline experience is a plus Inject functional and true DevOps experience CI/CD pipeline function experience...Devops
- TypeContractPlatform Engineer - Neo4jLocation: RemoteWe are seeking a Platform Engineer with... ...leader while contributing to the reliability, security, and scalability of our platform... ...reliability and efficiencyProvide platform and DevOps support for a graph database application...Devops
$149.4k - $180k
OverviewJob PurposeThe Kubernetes Platform Engineering (KPE) team builds and operates ICE's internal container orchestration platform powered... ...and Experience6+ years of experience in platform engineering, DevOps, or a related disciplineHands-on experience with Kubernetes in...DevopsFull time- ...Job Description: We are seeking an experienced Software Engineer with DevOps expertise to join our Algorithm Development team. This role... ...essential functions of this position.When travelling to Client sites, essential requirements of this position may require physical...DevopsFull timeWork at office
- ...and cloud computing.Roles and Responsibilities Lead development efforts for Java based cloud native applications.Collaborate with DevOps teams to deploy and manage applications in cloud environments.Support migration activities from legacy systems to cloud native architectures...Devops
- ...Reference26-00314TypeContractExperience7Remote60% Remote Software Engineer We are seeking a highly motivated, forward-thinking... ...SCM is a must (experience with git is a plus). Understanding of DevOps processes. Strong organizational and troubleshooting skills with...DevopsWork experience placementRemote work
- ...Consultancy and Information Technology Enabled Services.Job DescriptionSCM System EngineerSCM Continuous Integration / Delivery Build Team Engineer with experience in Application Service and Web Application Build, Deployment and Release Management and experience in establishing...DevopsPermanent employmentFull timeH1b
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to DevOps / Site Reliability Engineer. Be the first to apply!
Related searches
- big data devops engineer Atlanta, GA
- devops cloud engineer Atlanta, GA
- devops engineer sre Atlanta, GA
- devops engineer remote Atlanta, GA
- senior devops engineer remote Atlanta, GA
- senior devops engineer Atlanta, GA
- devops engineer Atlanta, GA
- senior devops cloud engineer Atlanta, GA
- devops engineer full time Atlanta, GA
- devops aws developer (remote) Atlanta, GA



