Site Reliability Engineer
Black Knight Financial Services
OverviewJob PurposeAt Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect companies around the world to global capital and derivative markets. With a leading-edge approach to developing technology platforms, we have built market infrastructure in all major trading centers, offering customers the ability to manage risk and make informed decisions globally. By leveraging our core strengths in technology, we continue to identify new ways to serve our customers and transform global markets. We're looking for motivated, results-oriented people to join our team.We are seeking a Site Reliability Engineer II to bring 3+ years of hands-on experience to our SRE team, operating with significant autonomy to improve platform reliability, drive automation, and mentor junior engineers. The ideal candidate contributes meaningfully to platform release cycles, leads smaller projects, and actively shapes the team's approach to observability, incident response, and service design in ICE's 24x7 production environment.ResponsibilitiesEmploy advanced troubleshooting and root-cause analysis to improve availability, performance, and security of IMT and platform servicesCollaborate with Product and Engineering teams to plan and deploy product releases with operational rigor and quality gatesWork with Engineering leadership to build and evolve shared services meeting the requirements of platform and application teamsDesign and implement proactive monitoring, alerting, trend analysis, and self-healing automationResolve product and service defects, infrastructure issues, and operational changes with increasing independenceImplement automated tests, automated deployments, and operational tooling across the SRE toolchainEnsure services are designed with 24x7 availability and operational readiness and rigorLead smaller projects and provide status updates to management and stakeholdersMentor SRE I engineers and contribute actively to team training and knowledge-sharingPartner with application and platform teams to identify critical workflows and build automated health checks that run post-deployment and during incidents to accelerate root-cause identificationDesign and build AI-assisted automated diagnosis jobs that correlate signals across monitoring and alerting platforms to reduce Mean Time to Resolution (MTTR) for production incidentsBuild and maintain automation pipelines (e.g., Rundeck, Jenkins) that integrate with AI/LLM tooling to drive efficiency gains in observability, runbook execution, and incident triageDevelop and tune AWS CloudWatch metrics, alarms, and dashboards, instrument services using OpenTelemetry/Alloy, and build observability visualizations in Grafana; integrate alerting and event correlation workflows across PagerDuty, BigPanda, and Splunk to ensure timely, actionable incident notificationKnowledge and ExperienceBachelor's degree in Computer Science, Engineering, or equivalent experience3+ years of experience in a site reliability, production engineering, or software operations roleProven technical skills with strong personal initiative and consistent delivery of important workExcellent teamwork with active involvement in training and mentoringAbility to prioritize and execute without direct management guidanceStrong understanding of ICE Core CompetenciesPreferred Knowledge and ExperienceExperience in financial services technology, mortgage platforms, or exchange infrastructureFamiliarity with SRE principles including SLI, SLO, and error budget managementExposure to Terraform, Chef, Ansible, or equivalent infrastructure automation frameworksHands-on experience with AWS observability services, CloudWatch, Grafana, OpenTelemetry/Alloy, Splunk, BigPanda, PagerDuty, and job orchestration/automation platforms such as Rundeck and JenkinsPractical experience integrating AI/LLM-based tooling into operational workflows to automate diagnosis, reduce manual triage, and improve incident response efficiency#LI-JM1 Intercontinental Exchange, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to legally protected characteristics.
$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...SuggestedFull timeTemporary workWork experience placementFlexible hours- Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence...SuggestedWorldwide
$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SuggestedFull timeWork at officeLocal areaRemote workWork from home$98k - $148.5k
...Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office,you'll help build and operate the foundational infrastructure that...SuggestedWork at officeLocal areaFlexible hours- ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises...SuggestedPermanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$70 - $85 per hour
...redefine what’s possible, give shape to the future—and get there.What You’ll Do* Define and establish enterprise reliability standards, Site Reliability Engineering (SRE) practices, SLIs/SLOs, operational governance models, and engineering guardrails that enable scalable,...Temporary workLocal areaFlexible hours3 days per week- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...Temporary work
- ...out new infrastructure capabilities to improve platform reliability and scalability. Monitor system health using metrics,... .... Requirements At least 3 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles. Hands-on experience...Full timeWork at officeLocal areaFlexible hours
$167.7k - $245.2k
...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze...Full timeTemporary workLocal areaFlexible hours$111.61k - $131.3k
At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life,...Full timeWork experience placementLocal area3 days per week- ...our company effectively. The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's... ...infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline...Remote workFlexible hours
- ...Job Title: Site Reliability Engineer II (SRE II) Data & Intelligence Location: Atlanta, GA Contract Job Summary The Site Reliability Engineer II (SRE II) is responsible for ensuring the reliability, scalability, performance, security, and operational...Contract work
$151k - $297k
...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB's cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and...Local areaRemote workWorldwideFlexible hours- #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...
- ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's... ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps...Work at officeLocal areaRemote workWork from homeWorldwide
- ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it...Full timeLive inWork at office
- ...Consultancy and Information Technology Enabled Services.Job DescriptionSCM System EngineerSCM Continuous Integration / Delivery Build Team Engineer with experience in Application Service and Web Application Build, Deployment and Release Management and experience in establishing...Permanent employmentFull timeH1b
- ...technologies to enable scalable, secure, and reliable business operations. Applies strong... ...infrastructure.3. Manages infrastructure engineering projects and processes aligned with... ...benefit plans, please visit our Benefits site. Depending on the position and division,...Permanent employmentFull timePart timeWork experience placementH1bRemote workWork visaShift workWeekend workDay shift
$101.5k - $169.1k
...include an incentive program.Job DescriptionThe Release Train Engineer (RTE) has a primary purpose of supporting an Agile Release Train... ...organizational AI policies and standards. Monitor AI tool reliability across teams. Create backup plans for system failures. Maintain...Full timeWork at officeRemote workVisa sponsorshipFlexible hours- ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,... ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$141.3k - $237.4k
...AT&T, you won’t just imagine the future, you’ll build it.We are seeking a highly skilled and hands-on Lead Software Engineer to join Software Reliability Engineering (SRE) Onboarding and automation team. This role will drive innovation through automation, enhancement of...Full timeTemporary workWork at officeLocal areaRelocation$105k - $130k
...provide the high-speed capabilities our nation and its allies need to maintain a durable, asymmetric advantage. The Mission Systems Engineering (MSE) Team develops the Mission Management System (MMS)—a software platform that integrates mission subsystems, autonomy services...Weekly payPermanent employmentFull timeWork at office$68 - $75 per hour
DescriptionWe are seeking a Reliability Engineer for our customer at the CDC. In this role, the engineer will ensure the reliability, scalability, and operational health of EDAV’s Azure cloud environment. Terraform is central to this position: the engineer will independently...Contract workTemporary work$144k - $191k
...customers. We are looking for software engineers, hardware engineers, roboticists, and front... ...-in-the-loop demonstrations at test sites.Define a strategy to comply with appropriate... ...(V&V) plans to ensure robust and reliable system performance.Experience writing testable...Full timeWork experience placementWork at officeImmediate start$168.5k - $252.7k
...secure.About the RoleAs a Senior Software Engineer, you will play a key role in designing... ...mentor team members to ensure high-velocity, reliable delivery.What You’ll DoDesign, build, and... ...not Workday Careers. Please be aware of sites that may ask for you to input your data...Full timeContract workWork at officeRemote workHome officeFlexible hours- ...That’s how we’re UNSTOPPABLE for our employees!Are you ready for the next chapter in your Uncarrier journey? The Sr. System Reliability Engineer (SRE) guides and mentors other SREs and improves and protects the software and systems behind all of T-Mobile's IT services,...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours
- ...Reference26-00225 Job Title: ( Senior Software Configuration/Release Engineer ) About Kyyba: Founded in 1998 and headquartered in Farmington... ...structure combined with career development. Job Description On-Site Interviews Only HYBRID - IN THE OFFICE 2 DAYS PER WEEK Minimum...Work at officeVisa sponsorshipWork visa2 days per week
$149.4k - $180k
OverviewJob PurposeThe Lead Systems Engineer joins our Secrets and Vault Engineering team within Identity and Access Management. The team... ..., security, compliance) to translate requirements into reliable, well-governed services.Help shape the team's roadmap in emerging...Full timeImmediate start$165k - $190k
OverviewJob PurposeThe Kubernetes Platform Engineering (KPE) team builds and operates ICE's internal container orchestration platform powered by Red Hat OpenShift. KPE Features, the Solutions Engineering sub-team, serves as the primary interface between the platform and...Full time$101.6k - $152.4k
We are seeking a talented Senior Engineer I, Digital Solutions to join our team and take charge of designing, developing, and deploying... ...alarms, and reports.Travel: Willingness to travel to customer sites as required. Travel is roughly expected to be around 25% but is...Full timeTemporary workImmediate startRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- site recruiter Atlanta, GA
- junior website developer Atlanta, GA
- on site coordinator Atlanta, GA
- construction site safety Atlanta, GA
- site services specialist Atlanta, GA
- website content developer Atlanta, GA
- website coordinator Atlanta, GA
- on-site clinical research associate (traveling/remote) Atlanta, GA



