DevOps / Site Reliability Engineer
AgileEngine
Job Description
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards. WHY JOIN US If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment. WHAT YOU WILL DO - Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP). - Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations. - Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment. - Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads. - Serve as Incident Commander on major and critical incidents - running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure. - Own the post-incident loop - track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups. - Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle. - Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution. MUST HAVES - You must be authorized to work for ANY employer in the US (e.g., Green card holders, TN visa holders, GC EAD, H4 EAD, U4U with EAD), as we are unable to sponsor or take over employment visa sponsorship at this time; - 5+ years of experience . - In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles . - Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting . - Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment. - Proven track record of remediation follow-up - coordinating with teams and holding owners accountable until issues are fully closed. - Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences. - Direct experience authoring divisional/group incident-management playbooks and escalation procedures. - Fully autonomous. - Drives the architecture of complex automated runbooks and mentors Middle-level SREs. - Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz . - Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2) . - Upper-intermediate English level. NICE TO HAVES - PagerDuty - hands-on experience with on-call scheduling, alert routing, and incident orchestration. - ServiceNow - familiarity with incident, problem, and change management workflows and reporting. PERKS AND BENEFITS - Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget - Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews - Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm - Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands - Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized - Well-being & support : access local well-being programs and people-focused support tailored to your location
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards. WHY JOIN US If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment. WHAT YOU WILL DO - Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP). - Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations. - Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment. - Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads. - Serve as Incident Commander on major and critical incidents - running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure. - Own the post-incident loop - track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups. - Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle. - Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution. MUST HAVES - You must be authorized to work for ANY employer in the US (e.g., Green card holders, TN visa holders, GC EAD, H4 EAD, U4U with EAD), as we are unable to sponsor or take over employment visa sponsorship at this time; - 5+ years of experience . - In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles . - Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting . - Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment. - Proven track record of remediation follow-up - coordinating with teams and holding owners accountable until issues are fully closed. - Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences. - Direct experience authoring divisional/group incident-management playbooks and escalation procedures. - Fully autonomous. - Drives the architecture of complex automated runbooks and mentors Middle-level SREs. - Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz . - Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2) . - Upper-intermediate English level. NICE TO HAVES - PagerDuty - hands-on experience with on-call scheduling, alert routing, and incident orchestration. - ServiceNow - familiarity with incident, problem, and change management workflows and reporting. PERKS AND BENEFITS - Growth without limits : build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget - Competitive compensation : get recognition that reflects your skills and impact, with regular performance and compensation reviews - Flexibility : work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm - Meaningful, modern projects : build impactful products using modern technologies alongside global teams and leading brands - Collaborative culture : join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized - Well-being & support : access local well-being programs and people-focused support tailored to your location
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the DevOps / Site Reliability Engineer in Austin, TX vacancy
- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has... ...guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence... ...years in Site Reliability Engineering, DevOps, or Infrastructure Engineering with...DevopsWork at officeLocal area
- ...for a Senior SRE to join our Platform Engineering team as the operations owner of our observability... .... You’ll be responsible for the reliability, scalability, and continued evolution... ...experience 5+ years of experience in SRE, DevOps, or platform engineering roles Deep...DevopsFull time
- ...you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability... ..., IT, or equivalent practical experience.5+ years in SRE, DevOps, or software development roles.Strong experience with...DevopsTemporary workCasual workWorldwide
$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...DevopsFull timeWork at office- ...the rapidly evolving state of the art to engineer scalable, innovative, and research... ...developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems,... ...experience5+ years of experience in SRE, DevOps, platform engineering, or developer tooling...DevopsFull timeLocal area
$167.18k - $203.61k
...DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate... ...-based solutions to drive improvement in SRE/DevOps best practices and tooling. Play a role in...DevopsFull timeWork at officeRemote workFlexible hours- ...Job Description Job Description Senior Site Reliability Engineer - Developer Productivity & Tooling Location: Austin, TX Area (Remote... ...planning. Requirements ~5+ years of experience in SRE, DevOps, Platform Engineering, or Developer Productivity. ~...DevopsLocal areaRemote work
$152k - $195k
...Senior Site Reliability Engineer Austin, TX (Hybrid) SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million... ...infrastructure. Required Qualifications ~6+ years in SRE, DevOps, or Infrastructure roles, with significant production...Devops- ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce... ...zones. This role requires deep experience with AWS, DevOps practices, and automation to support and improve our complex...DevopsFull timeWork at office
$110.7k - $171.8k
...Participation in oncall rotation as a platform reliability escalation point Incident response... ...requirements. Collaborate with engineering teams across the organization to... ...~ Experience in learning and applying DevOps and SRE best practices. ~ Experience...DevopsWork experience placementWork at officeLocal area- ...join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking... ...reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform....DevopsRemote workFlexible hoursNight shift
- ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of... ...experience in Site Reliability Engineering, Platform Engineering, DevOps, Cloud Infrastructure, or Software Engineering. ~ Proven...DevopsFull timeWork at office
$131.6k - $210.3k
...world. Progress starts with you. Job Description We are seeking a highly skilled and experienced Staff Site Reliability Engineer to join our DevOps squad. This role will focus on leading technical initiatives, optimizing CI/CD pipelines, automating infrastructure...DevopsWork experience placementWork at officeLocal areaRemote work- ...selected candidate for this role to work on site in the specified location(s).We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve... ...Cloud Architect or Professional Cloud DevOps Engineer.Experience operating Confluent for Kubernetes...DevopsFull timeWork at office
- ...submitted with LinkedIn Profile*** SRE Engineer Must-haves: Candidate should have... ...should have proficiency in CI/CD & DevOps Tools Experience in SQL & Database Reporting... ...We are looking for a highly skilled Site Reliability Engineer (SRE) to design| build| and...Devops
- ...SRE DevOps EngineerTekfortune is a fast-growing consulting firm specialized in permanent, contract & project-based staffing services... ...experts can help you find the best job for you. Role: SRE DevOps Engineer Location: Austin TX Duration: 6 months Required Skills:...DevopsPermanent employmentContract workRemote work
- ...Description Job Description SRE Support Engineer - Observability While this position is... ...Slack and tickets, improving monitoring reliability, and reducing incident impact through... ...Technical Support Engineering, SRE support, DevOps, Platform Support, or similar ~...DevopsRemote work
- Excelon Solutions in Austin, TX seeks a Site Reliability Engineer / Database Reliability Engineer to design, deploy, and maintain CockroachDB clusters... ...backup, DR, and incident response; collaborate with DevOps to improve reliability and operational excellence. On-call...Devops
$112.11k - $190.66k
...Position Summary We are looking for an experienced SRE - Site Reliability Engineer to work with our North American Team. Your responsibility... ...with BGP and Anycast routing is a plus ● Experience with DevOps principles and concepts such as Infrastructure as Code (...DevopsFull timeWork at officeLocal areaMonday to Friday$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...Full time$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours$165k - $241.4k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing...
- ...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week
- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical...Full timeWork at office
$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours- ...Configuring permissions and roles ~3+ years of hands-on experience with CI/CD tools such as: ~ Jenkins, TeamCity, Octopus, or Azure DevOps Nice to Have: ~3+ years of experience participating in Agile/Scrum development environments ~ Prior experience working...DevopsFull timePart timeRemote work
- TransPerfect Is More Than Just a Job…Our greatest asset is our people, and nothing is more important to us than ensuring that everyone knows that. Each of our 100+ offices has its own individual identity, and each also has its own unique rewards.SummaryLocation: US-Austin...DevopsFull time
- ...delivery robots. Our team ensures seamless network connectivity, reliable fleet-wide releases with observability, and end-to-end ride-log... ...NixOS (the latter is a significant plus). Familiarity with DevOps practices, including CI/CD processes on GitHub, AWS, and...DevopsFull timeRemote workRelocation
- ...observability systems to monitor infrastructure performance and ensure reliability.Collaborate with cross-functional teams to refine... ...of 3 years, preferably 5 years, of experience in platform engineering or DevOps.Proficiency in Python and Bash for scripting and automation...Devops
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to DevOps / Site Reliability Engineer. Be the first to apply!
Related searches
- big data devops engineer Austin, TX
- devops engineer sre Austin, TX
- devops engineer remote Austin, TX
- senior devops engineer remote Austin, TX
- senior devops engineer Austin, TX
- devops engineer Austin, TX
- senior devops cloud engineer Austin, TX
- devops engineer full time Austin, TX
- devops engineer contract Austin, TX
- devops aws developer (remote) Austin, TX



