Senior Site Reliability Engineer
Plenful
Senior Site Reliability Engineer (SRE)
Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow.
This role is centered on operating real systems at scale — not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time.
You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving — from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.
What You'll Do
Reliability Engineering & System Ownership
- Define and implement SLIs, SLOs, and error budgets across core services.
- Own production system health: uptime, latency, and availability targets.
- Improve system resilience through proactive reliability work.
- Find and mitigate single points of failure across distributed systems.
Production Operations & Incident Response
- Take part in and improve on-call rotations and incident response.
- Lead incident triage, mitigation, and resolution in real time.
- Run blameless postmortems and follow through on action items.
- Build tooling and automation to cut MTTR (Mean Time to Recovery).
Observability & System Insight
- Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
- Improve signal quality to cut noise and alert fatigue.
- Build dashboards and alerts that reflect real system health and user impact.
- Use observability data to drive performance and reliability improvements.
Performance & Scalability
- Analyze system performance under load and find bottlenecks.
- Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
- Partner with engineering teams to improve system efficiency and scaling behavior.
Automation & Reliability Tooling
- Build automation that eliminates repetitive operational work.
- Improve deployment safety through reliability checks and safeguards.
- Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
- Build tools for incident response, debugging, and capacity planning.
Security, Compliance & Operational Maturity
- Partner with security and compliance to keep systems meeting operational standards.
- Support audit readiness and reliability-related compliance requirements (Vanta).
- Integrate monitoring and alerting into security and SIEM workflows.
- Help mature operational practices across engineering.
You'll know it's working when SLOs and error budgets are clear and enforced, incidents are rare and shrinking over time, engineers trust their signals about system health, alerts are actionable instead of noisy, systems scale predictably under load, postmortems drive real improvement, and reliability is a shared responsibility, not a reactive function.
You May Be a Fit If
- You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
- You've operated and debugged distributed systems in production.
- You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
- You've defined and worked with SLOs, SLIs, and error budgets.
- You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
- You can write code or scripts (Python, Bash, etc.) for automation and tooling.
- You think in systems and reason clearly about failure modes.
- Bonus points for experience in high-growth or high-scale environments, background in regulated industries like healthcare or fintech, experience with ClickHouse or analytical systems at scale, familiarity with chaos engineering or load testing, and exposure to ML infrastructure or data platforms.
Why You'll Love Working Here
- Mission-Driven, World-Class Team — Join an exceptional group of professionals aligned around a meaningful mission and committed to making an impact
- Opportunities for Growth — Strengthen your expertise through collaboration with experienced, high-performing leaders across the organization
- Flexible Hybrid Work Environment — We're remote-first, with meaningful office presence in San Francisco and New York. R&D roles follow a hybrid model, with two days per week in our San Francisco office
Benefits & Perks
- Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family
- 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
- Equity — Every full-time employee shares in our success
- Unlimited PTO — Take the time you need, when you need it
- Daily Lunch Stipend — $100/week to cover your midday meals
- Wellness Stipend — $100/month to support your health and well-being
- Commuter Benefits — $100/month for SF and NYC-based employees
- Parental Leave — Paid leave to support growing families
- ...The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also...SeniorFull timeWork experience placementRemote work
- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....SeniorLocal area
$182.8k - $247.3k
...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement$150k - $170k
...Senior Site Reliability Engineer – Zip CoJoin to apply for the Senior Site Reliability Engineer role at Zip CoAt Zip, we build cloud-native software applications that serve millions of customers and process billions of dollars in payments. We're looking for a seasoned...SeniorCasual workWork at officeRemote work$200k - $240k
...Senior Site Reliability Engineer Inside health systems, where every second can matter, operations are still spread across dozens of disconnected tools and platforms. Kontakt.io is changing that. We combine proprietary hardware, AI-powered intelligence, and deep...SeniorWork at office3 days per week$168k - $200k
...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,...Senior$191k - $226k
....S. — and using AI to scale that impact further and faster than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads...SeniorRemote workWork visaFlexible hours- ...contributes. No one coasts. If you're driven by impact, pace, and raising the bar. This is the place. The role As a Senior Site Reliability Engineer you'll join the founding SRE team at our new NYC engineering hub, sitting within Foundations. You'll own critical...SeniorWork at office
$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...SeniorWork experience placementWork at officeLocal areaRelocation3 days per week$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience Technologies...SeniorFull timeLocal areaRemote workRelocation packageFlexible hoursShift work$139k - $257.55k
...The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...SeniorFull timeTemporary workLocal areaRemote workWorldwide- ...Job Description A major financial services company in NYC is growing its team rapidly, and they are looking for a Senior DevOps Engineer / Site Reliability Engineer who can join. If you’re passionate about high-availability, reliability, automation, we’d be excited...Senior
$156k - $262k
...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to... ...missing link between LLMs and the real world. The Role: Senior Site Reliability Engineer ~ Managing Kubernetes clusters across multiple...SeniorFull timeTemporary workWork at officeImmediate startRemote work$104k - $178k
...advertising ecosystem, helping to build a better industry. Learn more at About the Role Role Overview You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The team is responsible for building and maintaining the...Senior- ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure...Senior
- ...Principal Site Reliability Engineer Location: New York, NY (Onsite) Job type: Contract Job Description: Job Requirements Must Have: - Site Reliability Engineering and system reliability optimization - GitLab CI/CD and HashiCorp Vault secrets management -...Contract work
$400k
...in financial markets, the organization combines innovation, engineering excellence, and data-driven insights to support complex trading operations worldwide. This opportunity is for a Senior Site Reliability Engineer to join a high-performance infrastructure...SeniorPermanent employmentWorldwide$500 per month
...accounts. Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to... ...significant impact, we encourage you to apply. Your Role: As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform...SeniorHome office$150k - $190k
DescriptionKforce has a client that is seeking a Senior Principal Software Engineer (Delivery & Architecture) in New York, NY.Overview:We are seeking... ...in ambiguous environments and prioritizes high-quality, reliable delivery.Key Responsibilities:* Oversee the end-to-end...Senior$207k - $300k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Master's degree in Computer Science or Engineering.Experience mentoring engineers and cultivating... ...consensus across cross-functional teams.Site Reliability Engineering (SRE) combines...$120k - $150k
...allows each person to achieve personal success and add value to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About the RoleJoin our Platform Engineering team as aSite Reliability...Full time- ...Plaid Inc is looking for a Staff Site Reliability Engineer to lead the reliability practices across product engineering. You will architect SLO and error-budget programs, ensuring new products are production-ready while promoting safety gates for high velocity. The...
$208.5k - $216.5k
...pipelines - including its signal quality and its cost.Drive reliability improvements using SLOs and telemetry data, closing observability... ...review. Typically 8+ years in SRE, DevOps or infrastructure engineering, though scope and impact weigh more than tenure.Technical...Full timeTemporary workFlexible hours$235k - $250k
...Growing. We are seeking a strategic, hands-on Global Manager of Site Reliability Engineering to lead the reliability, release engineering, and... ...engineering, product, architecture, security, and operations, the Senior Manager will ensure that reliability is engineered into...Permanent employmentFull timeLocal areaFlexible hours- ...a mutual company built to last.Role OverviewAs a Release Train Engineer at New York Life, you’ll enable an Agile Release Train (ART) to... ...about our comprehensive benefit options or visit our NYL Benefits Site.Our Commitment to InclusionAt New York Life, fostering an inclusive...SeniorLocal areaShift work3 days per week
- ...our clients to succeed in an evolving digital landscape. Role Overview We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will bridge the...
$160k - $240k
Senior Software Engineer - BQL Reliability Engineering Location New York Business Area Engineering and CTO Ref # 10053945 Description & Requirements What You’ll Do: As part of the BQL (Bloomberg Query Language) Reliability Engineering team, you...SeniorTemporary workFor contractorsWork experience placement$80k - $95k
...join our dynamic team supporting the company’s users, applications, and web-based product offerings. In this role, the Site Reliability Engineer (SRE) will play a key role in maintaining resources at peak efficiency to guarantee staff are able to perform their functions...Remote workVisa sponsorshipWork visa- ...Job Description Job Description Location: New York, NY, USA Exp: 8-12 Years Client: Amex Job Description: SRE Engineer (This is not a Devops role, strictly need an SRE Engineer, who has great analytical skills and is a good incident manager as well)...
$120k - $180k
...people, and works with high-profile manufacturers including leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield opportunity to architect the path from AWS to on-premises...Permanent employmentFull timeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- site reliability engineer sre New York, NY
- senior living director New York, NY
- heritage senior living New York, NY
- senior manager customer operations New York, NY
- senior support engineer New York, NY
- senior product manager mobile New York, NY
- senior c++ developer New York, NY
- senior java developer New York, NY



