Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$140k - $180k

Endear

Our Story

At Endear, we’re building a modern CRM for retail teams—starting with the frontline. Our software helps sales associates have more personal, effective customer conversations through AI-powered tools that drive measurable revenue.

Despite retail being significantly larger than eCommerce, most software overlooks the in-store experience. Endear helps brands turn real customer relationships into growth through intuitive software and thoughtful design.

As Endear grows, we are investing in the reliability and platform systems that keep our product fast, resilient, and easy for engineering teams to operate.

Position Overview

We’re hiring a Senior Site Reliability Engineer to become Endear’s first dedicated reliability hire. This is a hands-on builder role for someone who wants to solve the underlying systems problems that create on-call burden—not simply respond to pages.

You will own the work of making Endear’s systems more observable, reliable, and scalable. You’ll investigate recurring incidents, drive root-cause fixes, improve alerting and incident playbooks, and partner with engineers on database, queue, and event-processing reliability.

You will also help establish a nearshore triage layer for routine, well-understood issues, so product engineers can spend more time building product and less time on operational interruptions.

What You’ll Accomplish
In your first 6 months, you’ll…
  • Build a clear view of Endear’s highest-impact reliability risks, recurring incidents, and on-call pain points.

  • Establish a prioritized reliability backlog and drive root-cause fixes for the most important issues.

  • Improve alert quality, severity definitions, escalation paths, and runbooks for common incidents.

  • Strengthen observability, queue health, database capacity planning, and operational readiness ahead of peak retail periods.

  • Create the foundation for a nearshore triage process for low-priority, repeatable issues.

In your first year, you’ll…
  • Make on-call materially quieter and less disruptive for product engineers.

  • Build scalable systems for observability, alerting, incident response, database reliability, and queue/event-processing health.

  • Own and improve the nearshore triage relationship, playbooks, and escalation process.

  • Help establish the technical roadmap and future resourcing plan for Endear’s broader platform and reliability function.

You’ll Thrive in This Role If You…
  • Have deep hands-on experience with Kubernetes, production databases, and event-driven systems.

  • Have operated and improved high-volume production systems with meaningful reliability, performance, and data-scale requirements.

  • Enjoy finding root causes, fixing repeat incidents, and building tooling that makes engineers’ lives easier.

  • Have experience with observability, alerting, incident response, capacity planning, and operational runbooks.

  • Can work effectively as a senior IC: owning complex technical work directly while coordinating across teams.

  • Are comfortable in a lean environment where priorities move quickly and you will need to make practical trade-offs.

  • Bring experience from a scaling, mid-size company rather than only an early-stage startup or hyperscaler environment.

  • Have GCP or ClickHouse experience, which are strong pluses.

About the Team

You’ll partner with:

  • JP Grace, CTO: Align on reliability priorities, technical risks, and the roadmap for improving on-call health.

  • Engineering team: Partner on root-cause fixes, platform improvements, and architecture decisions that improve reliability.

  • Nearshore triage partner: Build and maintain playbooks, escalation paths, and expectations for routine incident handling.

  • Product and Support: Help ensure issues are surfaced, prioritized, and resolved with the right level of urgency.

Endear is a lean, remote team where individuals have broad ownership. This role will directly shape how the company handles production reliability as it grows.

Our Hiring Process
  1. Recruiter screen — 30 minutes

  2. Behavioral interview with CTO — 60 minutes

  3. Technical panel with Engineering — 60 minutes

  4. Final conversation with Co-Founders

  5. Offer

Compensation & Benefits
  • Base salary: $140,000-180,000

  • Fully remote, U.S.-based role

  • Comprehensive healthcare, including medical, dental, and vision, plus a 401(k) plan

  • Monthly stipend for co-working and home-office setup

  • Flexible PTO and unlimited vacation

  • Opportunity to build Endear’s first dedicated reliability function from the ground up

Apply Even If You Don’t Check Every Box

If this role excites you but you’re unsure if you meet every requirement, reach out anyway. We care about skills, motivation, and how you think more than perfect resumes.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating... 
    Senior
    Full time

    Ridgeline

    New York, NY
    11 hours ago
  •  ...love to meet you. Our Enterprise Information Technology (EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this role, you will move beyond traditional infrastructure... 
    Senior
    Permanent employment
    Full time
    H1b
    Local area
    Remote work
    Shift work

    Jack Henry & Associates

    New York, NY
    11 hours ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Senior
    Full time

    Vanguard

    Wayne, PA
    3 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    5 days ago
  •  ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will... 
    Senior
    Remote work

    SRI Tech

    Plano, TX
    5 days ago
  • $170k - $220k

    Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating... 
    Senior

    Supio

    Seattle, WA
    5 days ago
  • $104.9k - $174.7k

    About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding cost optimization into infrastructure design, automation, monitoring, and operational processes. A FinOps SRE proactively identifies... 
    Senior
    Full time
    Local area

    RELX Group

    Boca Raton, FL
    5 days ago
  • IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Senior
    Work at office
    Immediate start

    IXL Learning

    Raleigh, NC
    7 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    3 days ago
  • $182.8k - $247.3k

     ...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed... 
    Senior
    Work experience placement

    Duolingo

    Pittsburgh, PA
    7 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    7 days ago
  • $152.6k - $191.5k

     ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...and continuous improvement.Position Summary: The Senior Site Reliability Engineer acts as an advanced senior individual... 
    Senior
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, MI
    4 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Cambridge, MA
    5 days ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Senior
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    6 days ago
  • We are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering... 
    Senior

    Robert Half

    San Francisco, CA
    5 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    5 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  •  ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The...  ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and... 
    Senior

    Black Knight Financial Services

    Jacksonville, FL
    4 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    7 days ago
  • $147k - $202.4k

     ...career-defining work. We're all in on this mission. If you are too, let's talk.Our company is seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company specializing in securing large-scale systems. This role is a blend of software... 
    Senior
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    Bellevue, WA
    6 days ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...operations or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    5 days ago
  • $91.7k - $163.7k

     ...potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud environment in both the commercial and government cloud. The... 
    Senior
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Eden Prairie, MN
    6 days ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    3 days ago
  • The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams... 
    Senior
    Full time
    Work at office
    Local area

    Castleton Commodities International

    Stamford, CT
    5 days ago
  • $80k - $140k

    Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and... 
    Senior
    Full time
    Flexible hours
    Shift work

    Royal Bank of Canada

    Minneapolis, MN
    6 days ago
  • $15k

     ...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    3 days ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!