Senior Site Reliability Engineer
Plenful
About Plenful Plenful is on a mission to transform healthcare operations from the inside out. Fresh off our $50M Series B and backed by Notable Capital, Bessemer Venture Partners, TQ Ventures, Susa/Kivu Ventures, and other leading investors, we're building the category-defining AI automation platform that healthcare teams rely on to operate smarter, faster, and more efficiently. Our technology empowers healthcare operators across hospital and health systems, pharmacies and payors to eliminate manual work, reduce administrative burden, and improve compliance, all while unlocking critical revenue to fund programs for their in-need patient populations. Built by healthcare operators for healthcare operators, Plenful is driven by a deep understanding of the challenges facing today's care teams. We're passionate about equipping healthcare workers with world-class tools that deliver real, measurable impact, and we're proud to serve 100+ leading health systems, pharmacies, and healthcare organizations across the country. If you're excited to help shape the future of healthcare, we'd love to meet you. Apply now to join our growing team.
About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale - not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time. You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving - from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid. What You'll Do Reliability Engineering & System Ownership
About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale - not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time. You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving - from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid. What You'll Do Reliability Engineering & System Ownership
- Define and implement SLIs, SLOs, and error budgets across core services.
- Own production system health: uptime, latency, and availability targets.
- Improve system resilience through proactive reliability work.
- Find and mitigate single points of failure across distributed systems.
- Take part in and improve on-call rotations and incident response.
- Lead incident triage, mitigation, and resolution in real time.
- Run blameless postmortems and follow through on action items.
- Build tooling and automation to cut MTTR (Mean Time to Recovery).
- Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
- Improve signal quality to cut noise and alert fatigue.
- Build dashboards and alerts that reflect real system health and user impact.
- Use observability data to drive performance and reliability improvements.
- Analyze system performance under load and find bottlenecks.
- Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
- Partner with engineering teams to improve system efficiency and scaling behavior.
- Build automation that eliminates repetitive operational work.
- Improve deployment safety through reliability checks and safeguards.
- Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
- Build tools for incident response, debugging, and capacity planning.
- Partner with security and compliance to keep systems meeting operational standards.
- Support audit readiness and reliability-related compliance requirements (Vanta).
- Integrate monitoring and alerting into security and SIEM workflows.
- Help mature operational practices across engineering.
- You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
- You've operated and debugged distributed systems in production.
- You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
- You've defined and worked with SLOs, SLIs, and error budgets.
- You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
- You can write code or scripts (Python, Bash, etc.) for automation and tooling.
- You think in systems and reason clearly about failure modes.
- Bonus points for experience in high-growth or high-scale environments, background in regulated industries like healthcare or fintech, experience with ClickHouse or analytical systems at scale, familiarity with chaos engineering or load testing, and exposure to ML infrastructure or data platforms.
- Mission-Driven, World-Class Team - Join an exceptional group of professionals aligned around a meaningful mission and committed to making an impact
- Opportunities for Growth - Strengthen your expertise through collaboration with experienced, high-performing leaders across the organization
- Flexible Hybrid Work Environment - We're remote-first, with meaningful office presence in San Francisco and New York. R&D roles follow a hybrid model, with two days per week in our San Francisco office
- Healthcare Coverage - Full medical, dental, and vision insurance for you and participation for your family
- 401(k) with Company Match - Plenful matches 50% of your first 3% contributed
- Equity - Every full-time employee shares in our success
- Unlimited PTO - Take the time you need, when you need it
- Daily Lunch Stipend - $100/week to cover your midday meals
- Wellness Stipend - $100/month to support your health and well-being
- Commuter Benefits - $100/month for SF and NYC-based employees
- Parental Leave - Paid leave to support growing families
Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
$250k
...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves...SeniorFull timeRemote work- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ...a pivotal role in engineering the reliable, globally connected, multi-cloud... ...Role OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a strong...SeniorLocal areaRemote workWorldwideFlexible hours$190.8k - $267.1k
...helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser... ...team partners closely with Ads Engineering to improve reliability, scalability, operational... ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,...SeniorFor contractorsWork experience placement$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...SeniorFlexible hours- ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ...ownership, working together to build scalable, reliable, and secure products that empower... ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...SeniorTemporary workLocal areaWorldwide
$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content... ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering...Senior$165k - $227k
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly...SeniorLocal areaWorldwideFlexible hours- ...Airbyte Infrastructure And Reliability Engineer Airbyte is the data and action layer for AI agents. We give agents fast, accurate, authenticated access to business data across hundreds of sources, so they can discover the entities that matter, reason over real-time...SeniorWork at officeLocal areaFlexible hours
- ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here,... ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...SeniorWorldwideWeekend work
$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...SeniorWork experience placementWork at officeLocal areaRelocation3 days per week$117k - $209.33k
...Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team...SeniorFull timeFor contractors- ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and...SeniorImmediate startRemote workWorldwide
- Job Title At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition...SeniorTemporary workWork experience placement
$181k - $225k
...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time...SeniorWork at officeImmediate start3 days per week$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...SeniorTemporary workImmediate startFlexible hoursShift work$181k - $263k
...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a...SeniorFull timeWork from homeWorldwideFlexible hoursNight shift$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...with self-service Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...SeniorPermanent employmentLocal areaWorldwideFlexible hours$139.76k - $287.75k
...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,... ...will be instrumental in advancing the reliability, scalability, automation, observability... ...The ideal candidate is a highly hands-on engineer with strong production experience and a...SeniorWork at officeLocal areaRelocationRelocation package$106k - $130k
..., for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and...SeniorHourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours$15k
...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorWork at officeLocal areaRemote work$232.34k - $290.42k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We're proud... ...integrate our teams, systems, and career sites. About the Role Fivetran and dbt...SeniorFull timeWork at officeRemote workFlexible hours$300k
...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...SeniorPermanent employment- ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems... ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes...Senior
- ...design of information and operational support systems. Required Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted. Proven working experience in installing, configuring, and troubleshooting...Full timeWork experience placementRemote workFlexible hours
$61k - $101k
...Salary: $61,000 - 101,000 per year Requirements: We expect formal training or certification in site reliability engineering, plus 3+ years of hands-on experience. We want strong familiarity with SRE culture and the practical application of reliability principles...Full time- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business...
- ...guarantees and certifications. We're hiring staff-level SREs to help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions for how we deploy, observe, and secure the platform, and help...Remote workFlexible hours
$113.4k - $162k
...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at...Temporary work$150k
...Job Description Job Description About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- senior living director San Francisco, CA
- senior php developer remote San Francisco, CA
- senior manager customer operations San Francisco, CA
- senior support engineer San Francisco, CA
- senior product manager mobile San Francisco, CA
- senior software engineer ruby on rails San Francisco, CA
- sr finance manager San Francisco, CA


