Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, Observability

$160k - $200k
Full-time

Ripple

At Ripple, we’re building a world where value moves like information does today. It’s big, it’s bold, and we’re already doing it. Through our crypto solutions for financial institutions, businesses, governments and developers, we are improving the global financial system and creating greater economic fairness and opportunity for more people, in more places around the world. And we get to do the best work of our career and grow our skills surrounded by colleagues who have our backs.

If you’re ready to see your impact and unlock incredible career growth opportunities, join us, and build real world value.

At Ripple, we’re building a world where value moves like information does today. Through our crypto solutions for financial institutions, businesses, governments, and developers, we are improving the global financial system and creating greater economic fairness and opportunity for more people, in more places around the world.

Ripple Treasury, now a Ripple solution acquired in 2025, marks a significant expansion into the multi-trillion-dollar corporate finance arena. With more than 40 years of experience supporting some of the world’s largest and most sophisticated companies, Ripple Treasury integrates a treasury command center into Ripple’s technology stack—giving corporates the ability to move, manage, and optimize liquidity in real-time, across traditional and digital assets, under one expanded umbrella.

THE WORK:

This is an engineering-first role with a coaching dimension—not the other way around. You will spend the majority of your time doing hands-on observability and reliability engineering work: building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems. Alongside that, you will coach and consult with stream-aligned product teams, helping them build operational maturity over time.

You will join Ripple’s Technical Operations team and work across Azure (80%) and AWS (20%) environments supporting infrastructure that is predominantly Windows-based (80%), handling significant payment volume for enterprise treasury customers. The incident management program you will help build is early-stage—you will be establishing practices, not inheriting a mature playbook.


WHAT YOU’LL DO:

  • Observability Engineering

    • Design and implement monitoring, alerting, and dashboards in New Relic (APM, Infrastructure, Logs, Synthetics) across Azure and AWS; write NRQL queries for troubleshooting, analysis, and reporting.
    • Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and communicate system health to stakeholders.
    • Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.
    • Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.
    • Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation (RED/USE methods), distributed tracing, and effective dashboard patterns.

    Infrastructure & IaC

    • Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources, alert configurations, and observability infrastructure—this is a primary engineering responsibility, not an occasional task.
    • Establish and enforce IaC governance standards for observability infrastructure across teams, providing a repeatable, auditable model for how monitoring resources are managed.
    • Author and troubleshoot Azure DevOps pipelines; support teams with deployment visibility, change tracking, and release hygiene as it relates to production reliability.

    Incident Management

    • Administer and configure Incident.IO: alert routing, notification workflows, Slack and OpsGenie integration, and runbook management—operationalizing what exists today and expanding from there.
    • Build out incident management foundations that are largely yours to establish: PIR/postmortem processes, on-call rotation design, escalation policies, incident severity classification, and response playbooks.
    • Track and report on MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement in partnership with engineering teams.
    • Respond to and debrief on production incidents—providing real-time troubleshooting support and facilitating structured post-incident reviews.

    Cross-Functional Enablement

    • Enable stream-aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands-on guidance.
    • Collaborate with the Subsystems Platform Team to translate common needs into self-service observability and incident management capabilities.
    • Build lasting team competency through documentation, training materials, and knowledge-sharing sessions that outlast any individual engagement.

WHAT YOU'LL BRING:

  • Core SRE Experience

    • 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations.
    • Proven ability to deliver hands-on engineering work while coaching and mentoring teams—comfortable switching between builder and consultant modes.
    • Experience working in Agile/Scrum environments and collaborating effectively with cross-functional teams.
    Observability & Incident Management Expertise — Required
    • Expert-level hands-on experience with New Relic (APM, Infrastructure, Logs, Synthetics, Alerts) and strong NRQL proficiency for troubleshooting and analysis.
    • Deep understanding of structured logging, metrics collection (RED/USE methods), distributed tracing, and designing effective dashboards and alerts.
    • Expertise defining and implementing SLOs/SLIs and error budgets for reliability management.
    • Hands-on experience with incident management platforms (Incident.IO, PagerDuty, OpsGenie, or similar).
    • Experience designing incident response workflows, on-call rotations, escalation policies, and facilitating post-incident reviews that drive actionable improvements.
    • Demonstrated ability to troubleshoot complex production issues using observability data across distributed systems.
    Infrastructure & Tools — Required
    • Strong Terraform experience: developing and maintaining IaC for cloud infrastructure and monitoring resources; familiarity with IaC governance patterns.
    • Proficiency with PowerShell scripting (required given the 80% Windows environment).
    • Strong experience with Azure cloud (App Services, Virtual Machines, Azure SQL, networking, monitoring) and working knowledge of AWS.
    • Experience with Azure DevOps for CI/CD pipeline authoring and troubleshooting.
    • Experience with Octopus Deploy for deployment management and release orchestration.
    • Comfort working across both Windows and Linux server environments.
    • Familiarity with Slack for operational workflows, alert routing, and incident communication.
    Desired / Additional
    • Experience with alert noise reduction strategies and observability cost optimization (log ingestion, pipeline rules, cardinality management).
    • Background facilitating chaos engineering, game day exercises, or failure injection to build team resilience.
    • Knowledge of VM-hosted SQL Server monitoring and performance optimization.
    • Familiarity with FinTech compliance requirements (SOC 2, ISO 27001) and audit evidence collection.
    • Experience measuring and improving key reliability metrics (MTTR, MTTD, availability, error budgets) at an organizational level.
    • Python or Bash scripting experience in addition to PowerShell.
    • Familiarity with Jira for incident tracking and workflow automation.
    Other common names for this role: Senior Site Reliability Engineer, Observability Engineer, Incident Management Engineer

For positions that will be based in NY, the annual salary range for this position is below. Actual salaries may vary based on numerous factors including, among other things, an individual applicant’s experience and qualifications for the position. This range does not include equity or additional compensation, such as bonuses or commissions.

NY Annual Base Salary Range

$160,000—$200,000 USD

WHO WE ARE:

Do Your Best Work

  • The opportunity to build in a fast-paced start-up environment with experienced industry leaders
  • A learning environment where you can dive deep into the latest technologies and make an impact. A professional development budget to support other modes of learning.
  • Thrive in an environment where no matter what race, ethnicity, gender, origin, or culture they identify with, every employee is a respected, valued, and empowered part of the team.
  • In-office collaboration for moments that matter is important to our culture, and we give managers and teams the flexibility to decide which 10+ days a month they come in.
  • Bi-weekly all-company meeting - business updates and ask me anything style discussion with our Leadership Team
  • We come together for moments that matter which include team offsites, team bonding activities, happy hours and more!

Take Control of Your Finances

  • Competitive salary, bonuses, and equity
  • Competitive benefits that cover physical and mental healthcare, retirement, family forming, and family support
  • Employee giving match
  • Mobile phone stipend

Take Care of Yourself

  • R&R days so you can rest and recharge
  • Generous wellness reimbursement and weekly onsite & virtual programming
  • Generous vacation policy - work with your manager to take time off when you need it
  • Industry-leading parental leave policies. Family planning benefits.
  • Catered lunches, fully-stocked kitchens with premium snacks/beverages, and plenty of fun events

Benefits listed above are for full-time employees.

Ripple is an Equal Opportunity Employer. We’re committed to building a diverse and inclusive team. We do not discriminate against qualified employees or applicants because of race, color, religion, gender identity, sex, sexual identity, pregnancy, national origin, ancestry, citizenship, age, marital status, physical disability, mental disability, medical condition, military status, or any other characteristic protected by local law or ordinance.

Please find our UK/EU Applicant Privacy Notice and our California Applicant Privacy Notice for reference.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, Observability in New York, NY vacancy
  •  ...Site Reliability Engineer Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute....  ...rotations...) • Experience working against reliability KPIs (observability, alerting, SLAs) • Hands-on experience with CI/CD,... 
    Suggested
    Relocation package

    Mistral AI

    New York, NY
    1 day ago
  • $104k - $178k

     ...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization....  ...be a hands-on technical contributor driving automation, observability, and operational excellence across the platforms that power... 
    Suggested

    DoubleVerify

    New York, NY
    3 days ago
  •  ...We are seeking a highly motivated Site Reliability Engineer (SRE) to join the Equity Trading Platform Engineering team, supporting critical...  ...connectivity, Linux/Unix and Windows environments, monitoring and observability tools, and automation. The role requires close... 
    Suggested
    Permanent employment
    Work at office
    Afternoon shift

    RIT Solutions, Inc.

    New York, NY
    18 hours ago
  • $120k - $150k

     ...teams and communities. We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY....  ...production systems and respond to incidents using enterprise observability tooling, contributing to alerting, dashboards, and SLO... 
    Suggested

    Piper Sandler Companies

    New York, NY
    4 days ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven...  ...work across cloud infrastructure, developer platforms, observability, and production workloads, including product applications... 
    Suggested
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    New York, NY
    3 days ago
  • $200k - $240k

     ...meet you! About the role We're looking for a Senior Site Reliability Engineer to join our Infrastructure Engineering team and get their...  ...knows exactly what happens when things break Build out observability that's genuinely tuned — SLIs, SLOs, and alerting people... 
    Work at office
    3 days per week

    Socket

    New York, NY
    2 days ago
  • $100k - $250k

     ...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial...  ...design, own, and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key... 
    Local area

    Kalshi

    New York, NY
    3 days ago
  • $260k - $300k

     ...makers of Devin, the first AI software engineer. Our team is extremely talent-dense....  ...expects. You will own both the production reliability of our user-facing products and the...  .... Build the monitoring, alerting, and observability systems that give the team a clear,... 

    Cognition AI

    New York, NY
    1 day ago
  • $120k - $180k

     ...in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield...  ...customer-specific cloud environments, and improve observability and delivery systems across the platform. Requirements... 
    Permanent employment
    Full time
    Relocation package

    Raydar Inc

    New York, NY
    3 days ago
  • $189k - $283.6k

     ...proactively and reactively improve the reliability of Block's platform and critical infrastructure...  ...tooling and automation to enhance observability, accelerate incident detection and...  ...strong desire to perform and grow as an engineer ~5+ years of software development experience... 
    Full time
    Local area
    Remote work
    Relocation package
    Flexible hours
    Shift work

    Block USA

    New York, NY
    3 days ago
  •  ...Software Reliability Engineer Good software has to run where customers need it. For many of Retool's largest customers, that means running...  ..., secret rotations, and migration steps. Improve observability for Retool Cloud, self-hosted customers, and internal operators... 

    re-tool®

    New York, NY
    2 days ago
  • $139k - $257.55k

     ...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,...  ...Languages: Python, PHP, bash, Ruby CI/CD: Jenkins, Argo CD Observability: New Relic, Splunk, Grafana, Prometheus About... 
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe

    New York, NY
    1 day ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and... 
    Flexible hours

    Baseten

    New York, NY
    3 days ago
  •  ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable...  ...and automation to cut MTTR (Mean Time to Recovery). Observability & System Insight Design and evolve observability across... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    New York, NY
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to...  ...troubleshooting complex trading infrastructure, managing observability, and collaborating across trading and technology teams to... 

    TheStaffed

    New York, NY
    4 days ago
  •  ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled...  ...implement preventive measures. Configure and manage observability solutions including logging, monitoring, tracing, and alerting... 
    Flexible hours

    Ova Technologies

    New York, NY
    18 hours ago
  • $130k - $200k

     ...real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads...  ...through to the retro. • Fluency with monitoring and observability; metrics, logs, dashboards, and alerting. • Comfort... 
    Shift work

    nScale

    New York, NY
    2 days ago
  •  ...Site Reliability Engineer Our Client, a multinational telecommunications technology company is seeking a Site Reliability Engineer (SRE...  ...performance of video services (live, linear, and on-demand) using observability tools. •Support day-to-day operations of video... 
    Temporary work

    Elite Technical

    New York, NY
    23 days ago
  • $191k - $226k

     ...anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of...  ...and rigorously reviewing infrastructure changes Own Observability: Build and maintain the monitoring, alerting, and observability... 
    Remote work
    Work visa
    Flexible hours

    Garner Health

    New York, NY
    23 days ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned...  ...orchestration infrastructure + Experience with observability of large scale distributed systems **About MongoDB** MongoDB... 
    Local area
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    18 hours ago
  • $500 per month

     ...team is a diverse group of experienced engineers, traders, and brokerage professionals who...  ...encourage you to apply. Your Role: As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform reliable, observable, and operable as we grow - working... 
    Home office

    Alpaca

    New York, NY
    24 days ago
  •  ...the bar. This is the place. The role As a Senior Site Reliability Engineer you'll join the founding SRE team at our new NYC engineering...  ...with full ownership Build and maintain a high-signal observability stack (metrics, logs, traces) and translate signals into... 
    Work at office

    Legora

    New York, NY
    3 days ago
  •  ...Job Title: Site Reliability Engineer (SRE) Job Location: New York, NY Job Type: Contract Job description : # Owning infrastructure...  ...runs reliably across multiple tenants # Building observability and monitoring to improve signal from existing... 
    Full time
    Contract work

    Staffingine LLC

    New York, NY
    1 day ago
  •  ...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment...  ...coordination when issues emerge. Enhance production observability and reliability by improving instrumentation, dashboards,... 
    Shift work
    New York, NY
    10 days ago
  • $61k - $101k

     ...year Requirements: Formal training or certification in site reliability engineering, plus 3+ years of applied experience Strong grasp of...  ...coordination when issues arise Improve production observability and reliability by enhancing instrumentation, dashboards... 
    Full time
    Shift work

    J.P. Morgan

    New York, NY
    4 days ago
  • $194k - $267k

     ...on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes...  ...provide service-to-service communication, security, and observability within the Kubernetes clusters. Enable fine-grained... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    New York, NY
    18 hours ago
  •  ...the bar. This is the place. The Role As a Staff Site Reliability Engineer you'll play a lead role on the founding SRE team at our new...  ...strategy across multiple teams and services Own the observability, capacity planning, and monitoring strategy for complex distributed... 
    Work at office

    Legora

    New York, NY
    3 days ago
  •  ...Staff Site Reliability Engineer Tabs is the leading AI-native revenue platform for modern finance and accounting teams. Tabs agents automate...  ...to design, build, and operate systems that are reliable, observable, and easy to develop on. You'll own our infrastructure... 
    Full time
    Contract work
    Work at office

    TABS inc.

    New York, NY
    3 days ago
  • We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep... 

    EPAM Systems

    New York, NY
    3 days ago
  •  ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized firm and have a direct and significant...  ...and root-cause practices, expanding SRE adoption (observability, monitoring, automation, operational analytics), and improving... 

    Hackajob

    New York, NY
    18 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, Observability. Be the first to apply!