Site Reliability Engineer
Intercontinental Exchange Holdings, Inc.
Site Reliability Engineer
At Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect companies around the world to global capital and derivative markets. With a leading-edge approach to developing technology platforms, we have built market infrastructure in all major trading centers, offering customers the ability to manage risk and make informed decisions globally. By leveraging our core strengths in technology, we continue to identify new ways to serve our customers and transform global markets. We're looking for motivated, results-oriented people to join our team.
We are seeking a Site Reliability Engineer to bring 3+ years of hands-on experience to our SRE team, operating with significant autonomy to improve platform reliability, drive automation, and mentor junior engineers. The ideal candidate contributes meaningfully to platform release cycles, leads smaller projects, and actively shapes the team's approach to observability, incident response, and service design in ICE's 24x7 production environment.
Responsibilities:
- Employ advanced troubleshooting and root-cause analysis to improve availability, performance, and security of IMT and platform services
- Collaborate with Product and Engineering teams to plan and deploy product releases with operational rigor and quality gates
- Work with Engineering leadership to build and evolve shared services meeting the requirements of platform and application teams
- Design and implement proactive monitoring, alerting, trend analysis, and self-healing automation
- Resolve product and service defects, infrastructure issues, and operational changes with increasing independence
- Implement automated tests, automated deployments, and operational tooling across the SRE toolchain
- Ensure services are designed with 24x7 availability and operational readiness and rigor
- Lead smaller projects and provide status updates to management and stakeholders
- Mentor SRE I engineers and contribute actively to team training and knowledge-sharing
- Partner with application and platform teams to identify critical workflows and build automated health checks that run post-deployment and during incidents to accelerate root-cause identification
- Design and build AI-assisted automated diagnosis jobs that correlate signals across monitoring and alerting platforms to reduce Mean Time to Resolution (MTTR) for production incidents
- Build and maintain automation pipelines (e.g., Rundeck, Jenkins) that integrate with AI/LLM tooling to drive efficiency gains in observability, runbook execution, and incident triage
- Develop and tune AWS CloudWatch metrics, alarms, and dashboards, instrument services using OpenTelemetry/Alloy, and build observability visualizations in Grafana; integrate alerting and event correlation workflows across PagerDuty, BigPanda, and Splunk to ensure timely, actionable incident notification
Knowledge and Experience:
- Bachelor's degree in Computer Science, Engineering, or equivalent experience
- 3+ years of experience in a site reliability, production engineering, or software operations role
- Proven technical skills with strong personal initiative and consistent delivery of important work
- Excellent teamwork with active involvement in training and mentoring
- Ability to prioritize and execute without direct management guidance
- Strong understanding of ICE Core Competencies
Preferred Knowledge and Experience:
- Experience in financial services technology, mortgage platforms, or exchange infrastructure
- Familiarity with SRE principles including SLI, SLO, and error budget management
- Exposure to Terraform, Chef, Ansible, or equivalent infrastructure automation frameworks
- Hands-on experience with AWS observability services, CloudWatch, Grafana, OpenTelemetry/Alloy, Splunk, BigPanda, PagerDuty, and job orchestration/automation platforms such as Rundeck and Jenkins
- Practical experience integrating AI/LLM-based tooling into operational workflows to automate diagnosis, reduce manual triage, and improve incident response efficiency
Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.
- ...solving and decision-making abilities and the highest degree of professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The infrastructure cloud team is responsible for internal services that provide...Suggested
$100k - $120k
...OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...SuggestedFull timeTemporary workWork experience placementFlexible hours- ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and...Suggested
$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SuggestedFull timeWork at officeLocal areaRemote workWork from home- ...Site Reliability Engineer Seek Now is transforming property inspections through technology, data, and human expertise. We deliver faster, smarter, more reliable insights to insurance carriers and single-family rental markets, and we're just getting started. If you want...SuggestedFlexible hours
$55 - $60 per hour
...Description Our client is looking for an SRE that will Lead the reliability, scalability, security, and operational excellence of... ...solutions to improve operational efficiency. • Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance...$120k - $175k
...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues...Full timeRemote workWork visaFlexible hours$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential and is located in Atlanta, GA (Onsite). Please review the job description below and contact me ASAP if you are interested...Contract workLocal areaImmediate start- ...and we're looking for new team members who want to be a part of this journey! We're looking for a proactive, hands-on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in fast-moving startup environments. You're someone who enjoys owning...Work experience placementFlexible hours
- ...We are currently looking for a Senior Software Engineer to be a part of the Site Reliability Engineering (SRE) team in Atlanta, GA . The SRE team is an innovative team devoted to providing a Docker-based Platform as a Service and assisting a growing number of teams...Contract workWork at officeLocal area
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Lead Engineer, Site Reliability Engineering Team As a lead engineer with Retail, Site Reliability Engineering team, you will be at the forefront of Cloud and Big Data technology. In this role you will establish yourself as a technical leader by exposing yourself to...
$121.4k - $218.6k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$71.6k - $119.4k
...support application teams. Our services provide applications with reliability, security, and better customer experiences. About the Job:... ...automation, troubleshoot issues, and work closely with senior engineers to learn and apply best practices. You’ll gain exposure to a...Full timeTemporary workInternshipLocal areaWork from home- ...availability. • Automation Experience with Build/deployment, Software Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in JavaEE/C++ Environments. • Experience in automating manual processes using Python, Ruby, Unix Shell (bash,...Immediate start
- ...Site Reliability Engineer (SRE) When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage...Work at office
$130k - $150k
...recruiter to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will...Full timeRemote work$123.4k - $222.53k
...! Ready grow your career as part of the Uncarrier journey at T-Mobile? Our team is searching for our next Sr. Site Reliability Engineer to strengthen the reliability and resilience of the systems powering T-Mobile's payment platforms, enabling faster, safer...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours$113k - $171.6k
...are growing rapidly and hiring top talent with leading AI skills across engineering, sales, product, marketing, and beyond as we build the leading digital operations platform. As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office,...Work at officeLocal areaFlexible hours- ...ideal time and number for communication, and the expected pay rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months Contract Visa : US Citizens/ Green Card Need Local to...Contract workLocal areaImmediate start
- ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively...Work experience placement
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...
- ...Technical Support Specialist In Site Reliability Engineering (Sre) Mandatory skills: Scripting and programming languages like Python, Java, Ruby. Cloud and infrastructure management – AWS, Google cloud and Azure is a plus- CI/CD Automation, Database Management. The...
- ...Site Reliability Engineer We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and...
- ...Senior Site Reliability Engineer Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems...Worldwide
$168k - $200k
...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...Remote work$178.13k - $205.4k
...Bachelor's degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5)... ...websites that are not Workday Careers. Please be aware of sites that may ask for you to input your data in connection with a job...Work at officeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer remote Atlanta, GA
- site reliability engineer sre Atlanta, GA
- site safety Atlanta, GA
- website coordinator Atlanta, GA
- on-site clinical research associate (traveling/remote) Atlanta, GA
- site services specialist Atlanta, GA
- on site coordinator Atlanta, GA
- construction site safety Atlanta, GA
- junior website developer Atlanta, GA

