Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, Observability

$160k - $200k
Full-time

Ripple

At Ripple, we’re building a world where value moves like information does today. It’s big, it’s bold, and we’re already doing it. Through our crypto solutions for financial institutions, businesses, governments and developers, we are improving the global financial system and creating greater economic fairness and opportunity for more people, in more places around the world. And we get to do the best work of our career and grow our skills surrounded by colleagues who have our backs.

If you’re ready to see your impact and unlock incredible career growth opportunities, join us, and build real world value.

At Ripple, we’re building a world where value moves like information does today. Through our crypto solutions for financial institutions, businesses, governments, and developers, we are improving the global financial system and creating greater economic fairness and opportunity for more people, in more places around the world.

Ripple Treasury, now a Ripple solution acquired in 2025, marks a significant expansion into the multi-trillion-dollar corporate finance arena. With more than 40 years of experience supporting some of the world’s largest and most sophisticated companies, Ripple Treasury integrates a treasury command center into Ripple’s technology stack—giving corporates the ability to move, manage, and optimize liquidity in real-time, across traditional and digital assets, under one expanded umbrella.

THE WORK:

This is an engineering-first role with a coaching dimension—not the other way around. You will spend the majority of your time doing hands-on observability and reliability engineering work: building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems. Alongside that, you will coach and consult with stream-aligned product teams, helping them build operational maturity over time.

You will join Ripple’s Technical Operations team and work across Azure (80%) and AWS (20%) environments supporting infrastructure that is predominantly Windows-based (80%), handling significant payment volume for enterprise treasury customers. The incident management program you will help build is early-stage—you will be establishing practices, not inheriting a mature playbook.


WHAT YOU’LL DO:

  • Observability Engineering

    • Design and implement monitoring, alerting, and dashboards in New Relic (APM, Infrastructure, Logs, Synthetics) across Azure and AWS; write NRQL queries for troubleshooting, analysis, and reporting.
    • Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and communicate system health to stakeholders.
    • Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.
    • Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.
    • Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation (RED/USE methods), distributed tracing, and effective dashboard patterns.

    Infrastructure & IaC

    • Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources, alert configurations, and observability infrastructure—this is a primary engineering responsibility, not an occasional task.
    • Establish and enforce IaC governance standards for observability infrastructure across teams, providing a repeatable, auditable model for how monitoring resources are managed.
    • Author and troubleshoot Azure DevOps pipelines; support teams with deployment visibility, change tracking, and release hygiene as it relates to production reliability.

    Incident Management

    • Administer and configure Incident.IO: alert routing, notification workflows, Slack and OpsGenie integration, and runbook management—operationalizing what exists today and expanding from there.
    • Build out incident management foundations that are largely yours to establish: PIR/postmortem processes, on-call rotation design, escalation policies, incident severity classification, and response playbooks.
    • Track and report on MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement in partnership with engineering teams.
    • Respond to and debrief on production incidents—providing real-time troubleshooting support and facilitating structured post-incident reviews.

    Cross-Functional Enablement

    • Enable stream-aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands-on guidance.
    • Collaborate with the Subsystems Platform Team to translate common needs into self-service observability and incident management capabilities.
    • Build lasting team competency through documentation, training materials, and knowledge-sharing sessions that outlast any individual engagement.

WHAT YOU'LL BRING:

  • Core SRE Experience

    • 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations.
    • Proven ability to deliver hands-on engineering work while coaching and mentoring teams—comfortable switching between builder and consultant modes.
    • Experience working in Agile/Scrum environments and collaborating effectively with cross-functional teams.
    Observability & Incident Management Expertise — Required
    • Expert-level hands-on experience with New Relic (APM, Infrastructure, Logs, Synthetics, Alerts) and strong NRQL proficiency for troubleshooting and analysis.
    • Deep understanding of structured logging, metrics collection (RED/USE methods), distributed tracing, and designing effective dashboards and alerts.
    • Expertise defining and implementing SLOs/SLIs and error budgets for reliability management.
    • Hands-on experience with incident management platforms (Incident.IO, PagerDuty, OpsGenie, or similar).
    • Experience designing incident response workflows, on-call rotations, escalation policies, and facilitating post-incident reviews that drive actionable improvements.
    • Demonstrated ability to troubleshoot complex production issues using observability data across distributed systems.
    Infrastructure & Tools — Required
    • Strong Terraform experience: developing and maintaining IaC for cloud infrastructure and monitoring resources; familiarity with IaC governance patterns.
    • Proficiency with PowerShell scripting (required given the 80% Windows environment).
    • Strong experience with Azure cloud (App Services, Virtual Machines, Azure SQL, networking, monitoring) and working knowledge of AWS.
    • Experience with Azure DevOps for CI/CD pipeline authoring and troubleshooting.
    • Experience with Octopus Deploy for deployment management and release orchestration.
    • Comfort working across both Windows and Linux server environments.
    • Familiarity with Slack for operational workflows, alert routing, and incident communication.
    Desired / Additional
    • Experience with alert noise reduction strategies and observability cost optimization (log ingestion, pipeline rules, cardinality management).
    • Background facilitating chaos engineering, game day exercises, or failure injection to build team resilience.
    • Knowledge of VM-hosted SQL Server monitoring and performance optimization.
    • Familiarity with FinTech compliance requirements (SOC 2, ISO 27001) and audit evidence collection.
    • Experience measuring and improving key reliability metrics (MTTR, MTTD, availability, error budgets) at an organizational level.
    • Python or Bash scripting experience in addition to PowerShell.
    • Familiarity with Jira for incident tracking and workflow automation.
    Other common names for this role: Senior Site Reliability Engineer, Observability Engineer, Incident Management Engineer

For positions that will be based in IL, the annual salary range for this position is below. Actual salaries may vary based on numerous factors including, among other things, an individual applicant’s experience and qualifications for the position. This range does not include equity or additional compensation, such as bonuses or commissions.

IL Annual Base Salary Range

$160,000—$200,000 USD

WHO WE ARE:

Do Your Best Work

  • The opportunity to build in a fast-paced start-up environment with experienced industry leaders
  • A learning environment where you can dive deep into the latest technologies and make an impact. A professional development budget to support other modes of learning.
  • Thrive in an environment where no matter what race, ethnicity, gender, origin, or culture they identify with, every employee is a respected, valued, and empowered part of the team.
  • In-office collaboration for moments that matter is important to our culture, and we give managers and teams the flexibility to decide which 10+ days a month they come in.
  • Bi-weekly all-company meeting - business updates and ask me anything style discussion with our Leadership Team
  • We come together for moments that matter which include team offsites, team bonding activities, happy hours and more!

Take Control of Your Finances

  • Competitive salary, bonuses, and equity
  • Competitive benefits that cover physical and mental healthcare, retirement, family forming, and family support
  • Employee giving match
  • Mobile phone stipend

Take Care of Yourself

  • R&R days so you can rest and recharge
  • Generous wellness reimbursement and weekly onsite & virtual programming
  • Generous vacation policy - work with your manager to take time off when you need it
  • Industry-leading parental leave policies. Family planning benefits.
  • Catered lunches, fully-stocked kitchens with premium snacks/beverages, and plenty of fun events

Benefits listed above are for full-time employees.

Ripple is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance.

Please find our UK/EU Applicant Privacy Notice and our California Applicant Privacy Notice for reference.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, Observability in Chicago, IL vacancy
  •  ...Observability / Site Reliability Engineer (SRE)Ontrac Solutions is a leading technology consulting firm, specializing in cutting-edge solutions that drive business transformation. We partner with organizations to modernize their infrastructure, streamline processes, and... 
    Suggested

    Ontrac Solutions Inc

    Chicago, IL
    11 hours ago
  • $250k - $350k

     ...where quantitative researchers, engineers, traders, and operational...  ...stability, throughput, and reliability Qualifications Minimum of 3...  ...experience in production support, site reliability, or...  ...production setting Familiarity with observability tools (e.g., Prometheus), and... 
    Suggested
    Full time

    Engtal Inc

    Chicago, IL
    2 days ago
  •  ...powered advice on this job and more exclusive features. Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency trading firm is seeking a mid to senior-level Site Reliability Engineer... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Algo Capital Group

    Chicago, IL
    2 days ago
  • $130k - $170k

     ...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end...  ...This position will lead the design and implementation of observability tools, incident response processes, and resilience strategies... 
    Suggested
    Full time
    Flexible hours
    Shift work

    Supernova Technology™

    Chicago, IL
    2 days ago
  • $152k - $205k

     ...Are you a systems-minded engineer who is happiest when production...  ...designed for? Do you want to own reliability for a platform that answers...  ...We’re looking for a Senior Site Reliability Engineer to join...  ...grow. You’ll work across observability, incident response, capacity... 
    Suggested
    Local area
    Remote work
    Work from home
    Visa sponsorship

    Fingerprint

    Chicago, IL
    2 days ago
  •  ...Play a key role in ensuring system reliability at one of the world's most iconic and...  ...largest financial institutions. As a Site Reliability Engineer II at JPMorgan Chase within the...  ...updating application code Understands observability patterns and strives to implement and... 
    Local area

    J.P. Morgan

    Chicago, IL
    4 days ago
  • $165k - $225k

     ...demanding AI workloads with enterprise-grade reliability and compliance. Your Role: You will...  ...core. Working closely with our systems engineers, network engineers, and platform...  ...reliability while establishing the automation, observability, and operational practices.  Job... 
    Remote work
    Flexible hours

    Moonlite

    Chicago, IL
    24 days ago
  • $136.2k - $214.01k

     ...Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the...  ...engineering teams to ensure Operations standards are observed, determine resource impacts for upcoming product... 
    Full time
    Flexible hours

    Proofpoint

    Chicago, IL
    4 days ago
  • $50 - $53 per hour

     ...Immediate need for a talented Site Reliability Engineer (SRE) This is a 12+ Months contract opportunity with long-term potential and is...  ...Function Apps, Logic Apps - Must ~8+ years Monitoring & Observability tools of which 3+ years with Grafana, Prometheus, KQL,... 
    Contract work
    Local area
    Immediate start

    Pyramid Consulting

    Chicago, IL
    2 days ago
  • $190.8k - $267.1k

     ...influential and trafficked corners of the internet. As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your...  ...will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for... 
    Work experience placement
    Home office
    Flexible hours

    Alien Blue

    Chicago, IL
    2 days ago
  • $130k - $180k

     ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and...  ...in on-call rotations. You’ll be a key voice in observability, change management, and service scalability, providing guidance... 
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to Friday
    Flexible hours

    iManage

    Chicago, IL
    a month ago
  • $108.08k - $172.5k

    Work with development and platform engineering teams to migrate and maintain applications in Google Cloud. Apply Observability concepts and applications to maintain services. Monitor metrics, system health and analyze reports. Provide on-call rotation support for production... 
    Remote work
    Worldwide

    CME Group

    Chicago, IL
    11 hours ago
  •  ...Edward Jones Site Reliability Engineer 100% remote Initial contract is 6 months, but will be a multi year engagement. Position Overview...  ...with development teams to integrate reliability and observability best practices into the software development lifecycle.... 
    Contract work
    Remote work

    HCL Global Systems

    Chicago, IL
    1 day ago
  •  ...Job Description As a Senior DevOps / SRE Engineer on contract, you will be embedded with the Central Technology AI enablement team...  ...or LangSmith Fleet, agent registries). · Experience with observability platforms (i.e. Langsmith) and cost-attribution patterns for... 
    Contract work
    Immediate start

    Insight Global

    Chicago, IL
    4 days ago
  •  ...tolerance for firms in active litigation. We're looking for a Site Reliability Engineer to help maintain the reliability, scalability, and...  ...Maintain and extend existing monitoring, alerting, and observability tooling Cross-Functional Collaboration (15%) Partner... 
    Work at office
    3 days per week

    Nextpoint

    Chicago, IL
    1 day ago
  • $160k - $210k

     ...you'll do:Join our Platform Engineering team, where you'll ensure the...  ...mentoring engineers across reliability initiativesAnalyze, troubleshoot...  ...of experience in DevOps, Site Reliability Engineering, or...  ...environmentsKnowledge of monitoring and observability tools such as Prometheus,... 
    Work at office
    Worldwide
    Monday to Friday
    Flexible hours

    NinjaTrader

    Chicago, IL
    2 days ago
  • $204k - $306k

     ...on this mission. If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from...  ...teams focused on Edge networking, K8s platform, CI/CD, Observability, automation platform & tooling. What you'll be doing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta, Inc.

    Chicago, IL
    1 day ago
  • $106.9k - $200.6k

     ...opportunity We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EY’s AI-native platform. These are the...  ...filtering, so every signal is captured and routed reliably. Automate GitOps-based delivery and... 
    Full time
    Summer holiday
    Flexible hours

    EY

    Chicago, IL
    2 days ago
  • $112.5k - $187.5k

     ...TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering...  ...targets. ~ Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog, Prometheus, Grafana... 
    Full time
    Work experience placement
    Work at office
    Flexible hours
    2 days per week

    TransUnion

    Chicago, IL
    1 day ago
  • $145k - $160k

     ...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep... 
    Temporary work
    Remote work
    Flexible hours

    EPAM Systems Inc

    Chicago, IL
    4 days ago
  • $194k - $267k

     ...-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes...  ...provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    Chicago, IL
    3 days ago
  • $150k - $170k

     ...global investing! About The Role As the Manager of Site Reliability Engineering, you'll lead a team of SRE Automation Engineers while...  ...in Python or Golang, along with Bash and Ansible. Observability : Experience with Grafana/Similar tools, Prometheus, Understanding... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    DriveWealth

    Chicago, IL
    1 day ago
  •  ...Hire Overview Our client is seeking a highly skilled Edge Site Reliability Engineer (Edge SRE) to lead the design, automation, and operations...  ...termination policies. Develop monitoring, alerting, and observability frameworks (real-time + historical) for latency, traffic,... 
    Full time
    Contract work

    CoSourcing Partners - Enterprise-AI and IT Services Company

    Chicago, IL
    2 days ago
  • Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources Willingness to work on-site at stated location in the job openingDepartment... 
    Contract work
    For contractors
    Work experience placement

    Cedent Consulting

    Chicago, IL
    3 days ago
  • $150k - $200k

     ...the job poster from Selby Jennings Recruitment Consultant @ Selby Jennings | Financial Technology We are seeking a Site Reliability Engineer to join our team and assist with the design, development, and administration of our trading and research systems. This role... 
    Full time
    Work at office

    Selby Jennings

    Chicago, IL
    2 days ago
  • $62 - $80 per hour

    Chicago, IllinoisRemote LocalContract$62/hr - $80/hrA senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant... 
    Full time
    Temporary work
    Remote work
    Flexible hours

    Motion Recruitment

    Chicago, IL
    2 days ago
  • $232k - $319k

     ...scale the service with great people and reliable, cost-effective, and efficient...  ...focused on Edge networking, K8s platform, Observability, automation platform & tooling.  What...  ...Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Chicago, IL
    19 days ago
  •  ...Reference: 26513Location: Chicago, IL, United StatesIndustry: Trading FirmPosted: Contact: Ethan HudsonEmail: : Job Title: Site Reliability Engineer (Infrastructure & Systems)Location: Chicago, IL (Greater Metro Area)About the OpportunityJoin a premier financial... 
    Local area

    Objective Paradigm

    Chicago, IL
    3 days ago
  • $95.1k - $122.55k

    About Us: At Clearwater, we are dedicated to provide world-class enterprise applications, ensuring their performance and availability to support our clients in the ever-evolving fintech landscape. Our Enterprise Application Support team plays a vital role in maintaining...
    Shift work

    Clearwater Analytics

    Chicago, IL
    1 day ago
  •  ...building and running systems that must perform reliably under real-time market conditions. The culture is highly collaborative, engineering-driven, and focused on continuous...  ...related field ~3+ years of experience in site reliability, systems engineering, or technical... 

    Fintal Partners

    Chicago, IL
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, Observability. Be the first to apply!