Site Reliability Engineer
Ice Services
Job Purpose
At Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect companies around the world to global capital and derivative markets. With a leading-edge approach to developing technology platforms, we have built market infrastructure in all major trading centers, offering customers the ability to manage risk and make informed decisions globally. By leveraging our core strengths in technology, we continue to identify new ways to serve our customers and transform global markets. We're looking for motivated, results-oriented people to join our team.
Overview
Job Purpose
At Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect companies around the world to global capital and derivative markets. With a leading-edge approach to developing technology platforms, we have built market infrastructure in all major trading centers, offering customers the ability to manage risk and make informed decisions globally. By leveraging our core strengths in technology, we continue to identify new ways to serve our customers and transform global markets. We're looking for motivated, results-oriented people to join our team.
We are seeking a Site Reliability Engineer to bring 3+ years of hands‑on experience to our SRE team, operating with significant autonomy to improve platform reliability, drive automation, and mentor junior engineers. The ideal candidate contributes meaningfully to platform release cycles, leads smaller projects, and actively shapes the team's approach to observability, incident response, and service design in ICE's 24x7 production environment.
Responsibilities
- Employ advanced troubleshooting and root-cause analysis to improve availability, performance, and security of IMT and platform services
- Collaborate with Product and Engineering teams to plan and deploy product releases with operational rigor and quality gates
- Work with Engineering leadership to build and evolve shared services meeting the requirements of platform and application teams
- Design and implement proactive monitoring, alerting, trend analysis, and self‑healing automation
- Resolve product and service defects, infrastructure issues, and operational changes with increasing independence
- Implement automated tests, automated deployments, and operational tooling across the SRE toolchain
- Ensure services are designed with 24x7 availability and operational readiness and rigor
- Lead smaller projects and provide status updates to management and stakeholders
- Mentor SRE I engineers and contribute actively to team training and knowledge-sharing
- Partner with application and platform teams to identify critical workflows and build automated health checks that run post‑deployment and during incidents to accelerate root‑cause identification
- Design and build AI‑assisted automated diagnosis jobs that correlate signals across monitoring and alerting platforms to reduce Mean Time to Resolution (MTTR) for production incidents
- Build and maintain automation pipelines (e.g., Rundeck, Jenkins) that integrate with AI/LLM tooling to drive efficiency gains in observability, runbook execution, and incident triage
- Develop and tune AWS CloudWatch metrics, alarms, and dashboards, instrument services using OpenTelemetry/Alloy, and build observability visualizations in Grafana; integrate alerting and event correlation workflows across PagerDuty, BigPanda, and Splunk to ensure timely, actionable incident notification
Knowledge And Experience
- Bachelor's degree in Computer Science, Engineering, or equivalent experience
- 3+ years of experience in a site reliability, production engineering, or software operations role
- Proven technical skills with strong personal initiative and consistent delivery of important work
- Excellent teamwork with active involvement in training and mentoring
- Ability to prioritize and execute without direct management guidance
- Strong understanding of ICE Core Competencies
Preferred Knowledge And Experience
- Experience in financial services technology, mortgage platforms, or exchange infrastructure
- Familiarity with SRE principles including SLI, SLO, and error budget management
- Exposure to Terraform, Chef, Ansible, or equivalent infrastructure automation frameworks
- Hands‑on experience with AWS observability services, CloudWatch, Grafana, OpenTelemetry/Alloy, Splunk, BigPanda, PagerDuty, and job orchestration/automation platforms such as Rundeck and Jenkins
- Practical experience integrating AI/LLM‑based tooling into operational workflows to automate diagnosis, reduce manual triage, and improve incident response efficiency
Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.
#J-18808-Ljbffr- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...SuggestedContract work
- ...in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.Site Reliability EngineerOnsite: Atlanta, GAJob SummaryAt NCR Voyix, we're looking for a Site Reliability Engineer II to help build, support, and scale the cloud platforms that power our...SuggestedFull timeWorldwideFlexible hours
- ...Site Reliability Engineer Full Description Company Overview: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12...SuggestedFull timeLive inWork at office
$100k - $120k
...Overview The Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...SuggestedFull timeTemporary workWork experience placementRemote workFlexible hours$178.13k - $205.4k
...partial telecommuting. Salary Range: $178,131 - $205,400 About You Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5) years of progressive, post-baccalaureate experience in the job...SuggestedWork at officeRemote workFlexible hours- ...data, and human expertise. We deliver faster, smarter, more reliable insights to insurance carriers and single-family rental... ...at scale, you’re in the right place. The Role As a Site Reliability Engineer, you'll be responsible for the availability, scalability,...Flexible hours
- ...new team members who want to be a part of this journey! Who We’re Looking For We’re looking for a proactive, hands‑on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in fast‑moving startup environments. You’re someone who enjoys...Work experience placementFlexible hours
$81.75k - $138.98k
...Locations 4125 GA hwy 316, Dacula, GA, 30019, US (Hybrid) Job Schedule Full time Job Description As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes...Full timeWork at officeImmediate startWorldwideShift work- ...focused on building and supporting advanced platforms and applications that drive our business forward. We are seeking a Site Reliability Engineer (SRE) to help define and raise the reliability bar for our Commerce Platform. As an SRE on the AI Commerce team, you will...Fixed term contract
- ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's... ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps...Work at officeLocal areaRemote workWork from homeWorldwide
- ...Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence...Worldwide
$130k - $145k
...Back Site Reliability Engineer Cloud/Infrastructure Atlanta , GA Sep 2, 2026 Site Reliability Engineer Atlanta, GA / Hybrid Blu Omega is seeking a Site Reliability Engineer to support a federal program focused on enterprise cloud modernization. This role operates...Temporary work$149.8k - $241.5k
...efficiency, accelerate time-to-value, and deliver better customer experiences. About The Role We're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'...Work at officeLocal areaRemote workWork from homeWorldwideHome officeFlexible hours$130k - $150k
...recruiter to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will...Full timeRemote work- ...Job Title :- Site Reliability Engineer (SRE) Employment Type :- W2 Duration :- Long Term Visa Type :- All Visa applicable which are ready for W2 Location :- Atlanta, GA (Onsite) Job Description We are seeking a highly skilled Site Reliability Engineer (SRE...
$141.8k - $195k
...best work, grow fast, and bring their full selves to the herd. Why You'll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl...Temporary workRemote work$123.4k - $222.53k
...Responsibilities Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime... ...) ~ Acceptable areas of study include Computer Science, Engineering or related field (Required) ~4-7 years Working in operations...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours- ...Join to apply for the Site Reliability Engineer role at Motion Recruitment Join to apply for the Site Reliability Engineer role at Motion Recruitment Get AI-powered advice on this job and more exclusive features. Every year, nearly 200 million travelers...Contract workWorldwide
- ...Site Reliability Engineer We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and...
$120k - $175k
...of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's possible...Full timeRemote workWork visaFlexible hours$113.2k - $188.8k
...to work on improving Grid resilience through Software, this is the job for you. We are looking for a Deployment and Site Reliability Engineer to join the GridBeats team. The GridBeats Software Portfolio aggregates all Grid Automation applications which monitor, optimize...Permanent employmentContract workRemote workRelocation- ...English (Required) Work Shift: 1st shift (United States of America) Please review the following job description: The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$61.09k - $104.36k
...breadth of their business needs. It delivers end-to-end services and solutions leveraging strengths from strategy and design to engineering, all fueled by its market leading capabilities in AI, generative AI, cloud and data, combined with its deep industry expertise and...Full timeLocal area- ...Site Reliability Engineering (SRE) Architect Location: Atlanta, GA Duration: 12Months+ Extension Hourly Rate: Depending on Experience (DOE) Work Authorization: As an SRE Architect, you will be a pivotal technical leader responsible for designing,...Hourly payPermanent employmentContract workLocal areaEarly shift
- ...our company effectively. The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's... ...infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline...Remote workFlexible hours
$70 - $75 per hour
...Site Reliability Engineering (SRE) Architect Get AI-powered advice on this job and more exclusive features. This range is provided by STAFFWORXS. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range...Contract work- ...Site Reliability Engineer We are looking for a Site Reliability Engineer to ensure the reliability, security, and continuous operation of a multi‑cloud application security platform. This role combines platform engineering and security automation, focusing on Kubernetes...Work at officeRemote workVisa sponsorshipWork visaFlexible hours
$178k - $213k
...Ventures, and Vista Credit Partners of Vista Equity Partners 2022 Cybersecurity Excellence Award for MDR Manager, Site Reliability Engineering Reports to: VP, Product Engineering Location: While proximity to Tampa is preferred to support hybrid schedule in...Permanent employmentWork experience placementWork at officeRemote workWork from homeHome officeFlexible hours- ...Role: Senior Site Reliability Engineer (SRE) Cloud & Kubernetes Location: Atlanta, GA (Onsite) Contract Role Summary: Lead the reliability, scalability, security, and operational excellence of customer-facing platforms across Azure, GCP, and Kubernetes...Contract work
$71.6k - $119.4k
...deployment support, and security improvements. You'll help implement automation, troubleshoot issues, and work closely with senior engineers to learn and apply best practices. You'll gain exposure to a wide range of cloud technologies, automation tools, and data...Temporary workInternshipLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer remote Atlanta, GA
- site reliability engineer sre Atlanta, GA
- site safety Atlanta, GA
- website coordinator Atlanta, GA
- on-site clinical research associate (traveling/remote) Atlanta, GA
- site services specialist Atlanta, GA
- on site coordinator Atlanta, GA
- construction site safety Atlanta, GA
- junior website developer Atlanta, GA




