Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Plenful

About Plenful

Plenful is on a mission to transform healthcare operations from the inside out. Fresh off our $50M Series B and backed by Notable Capital, Bessemer Venture Partners, TQ Ventures, Susa/Kivu Ventures, and other leading investors, we're building the category-defining AI workflow automation platform that healthcare teams rely on to operate smarter, faster, and more efficiently. Our technology empowers healthcare operators across hospital and health systems, pharmacies and payors to eliminate manual work, reduce administrative burden, and improve compliance, all while unlocking critical revenue to fund programs for their in-need patient populations.

Built by healthcare operators for healthcare operators, Plenful is driven by a deep understanding of the challenges facing today's care teams. We're passionate about equipping healthcare workers with world-class tools that deliver real, measurable impact, and we're proud to serve 90+ leading health systems across the country. If you're excited to help shape the future of healthcare, we'd love to meet you. Apply now to join our growing team.
About the Role

Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow.

This role is centered on operating real systems at scale - not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time.

You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving - from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.

What You'll Do

Reliability Engineering & System Ownership
  • Define and implement SLIs, SLOs, and error budgets across core services.
  • Own production system health: uptime, latency, and availability targets.
  • Improve system resilience through proactive reliability work.
  • Find and mitigate single points of failure across distributed systems.
Production Operations & Incident Response
  • Take part in and improve on-call rotations and incident response.
  • Lead incident triage, mitigation, and resolution in real time.
  • Run blameless postmortems and follow through on action items.
  • Build tooling and automation to cut MTTR (Mean Time to Recovery).
Observability & System Insight
  • Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
  • Improve signal quality to cut noise and alert fatigue.
  • Build dashboards and alerts that reflect real system health and user impact.
  • Use observability data to drive performance and reliability improvements.
Performance & Scalability
  • Analyze system performance under load and find bottlenecks.
  • Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
  • Partner with engineering teams to improve system efficiency and scaling behavior.
Automation & Reliability Tooling
  • Build automation that eliminates repetitive operational work.
  • Improve deployment safety through reliability checks and safeguards.
  • Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
  • Build tools for incident response, debugging, and capacity planning.
Security, Compliance & Operational Maturity
  • Partner with security and compliance to keep systems meeting operational standards.
  • Support audit readiness and reliability-related compliance requirements (Vanta).
  • Integrate monitoring and alerting into security and SIEM workflows.
  • Help mature operational practices across engineering.
You'll know it's working when SLOs and error budgets are clear and enforced, incidents are rare and shrinking over time, engineers trust their signals about system health, alerts are actionable instead of noisy, systems scale predictably under load, postmortems drive real improvement, and reliability is a shared responsibility, not a reactive function.

You May Be a Fit If
  • You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
  • You've operated and debugged distributed systems in production.
  • You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
  • You've defined and worked with SLOs, SLIs, and error budgets.
  • You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
  • You can write code or scripts (Python, Bash, etc.) for automation and tooling.
  • You think in systems and reason clearly about failure modes.
  • Bonus points for experience in high-growth or high-scale environments, background in regulated industries like healthcare or fintech, experience with ClickHouse or analytical systems at scale, familiarity with chaos engineering or load testing, and exposure to ML infrastructure or data platforms.
Why You'll Love Working Here
  • Mission-Driven, World-Class Team - Join an exceptional group of professionals aligned around a meaningful mission and committed to making an impact
  • Opportunities for Growth - Strengthen your expertise through collaboration with experienced, high-performing leaders across the organization
  • Flexible Hybrid Work Environment - We're remote-first, with meaningful office presence in San Francisco and New York. R&D roles follow a hybrid model, with two days per week in our San Francisco office
Benefits & Perks
  • Healthcare Coverage - Full medical, dental, and vision insurance for you and participation for your family
  • 401(k) with Company Match - Plenful matches 50% of your first 3% contributed
  • Equity - Every full-time employee shares in our success
  • Unlimited PTO - Take the time you need, when you need it
  • Daily Lunch Stipend - $100/week to cover your midday meals
  • Wellness Stipend - $100/month to support your health and well-being
  • Commuter Benefits - $100/month for SF and NYC-based employees
  • Parental Leave - Paid leave to support growing families
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    3 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • DescriptionWe are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and... 
    Senior

    Robert Half

    San Francisco, CA
    4 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    4 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    5 days ago
  •  ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ...ownership, working together to build scalable, reliable, and secure products that empower...  ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    3 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    2 days ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    2 days ago
  •  ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'... 
    Senior

    Alembic

    San Francisco, CA
    5 days ago
  •  ...founders with PhDs in AI, Math, and Computer Science - is poised to redefine computing. About the Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and... 
    Senior

    Hyperbolic Labs

    San Francisco, CA
    2 days ago
  • $160k - $250k

     ...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able... 
    Senior

    Hive

    San Francisco, CA
    2 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    3 days ago
  •  ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring through a focused Engineering Hiring Sprint...  ...: Platform Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers AI Agents Engineers... 
    Senior
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    3 days ago
  • $165k - $241.4k

     ...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    5 days ago
  •  ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it... 
    Senior
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    2 days ago
  • $210.8k - $272.8k

    Thumbtack is hiring for a Site Reliability Engineer to enhance the reliability and scalability of our services in San Francisco, CA. You'll design and support resilient systems, ensuring a smooth user experience. The ideal candidate will have extensive experience with AWS... 
    Senior

    Thumbtack

    San Francisco, CA
    3 days ago
  • $232k - $319k

     ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...  ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $250k

     ...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $175k - $250k

     ...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    4 days ago
  • $210.8k - $272.8k

    About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    3 days ago
  • CloudDevs works with fast-moving, venture-backed startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will either be placed directly into one of our partner startups or added to our... 
    Senior
    Local area

    Breakout Tools

    San Francisco, CA
    4 days ago
  • $15k

     ...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    1 day ago
  • $139.76k - $287.75k

     ...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,...  ...will be instrumental in advancing the reliability, scalability, automation, observability...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Senior
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We’re proud...  ...integrate our teams, systems, and career sites.About the RoleFivetran is building data... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    5 days ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We're proud...  ...integrate our teams, systems, and career sites. About the Role Fivetran is building... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    5 days ago
  •  ...better than we found it. The Apple Service Engineering (ASE) team builds and provides systems...  ...The ASE Compute team is looking for a senior SRE software engineer to own the...  ...management infrastructure, strengthen the reliability of our Kubernetes services, and engage... 
    Senior

    Socket.dev

    San Francisco, CA
    5 days ago
  • Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services...  ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role... 
    Senior

    Socket.dev

    San Francisco, CA
    5 days ago
  •  ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will... 
    Senior

    Kody

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!