Senior Site Reliability Engineer
Plenful
Senior Site Reliability Engineer (SRE)
Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow.
This role is centered on operating real systems at scale — not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time.
You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving — from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.
What You'll Do
Reliability Engineering & System Ownership
- Define and implement SLIs, SLOs, and error budgets across core services.
- Own production system health: uptime, latency, and availability targets.
- Improve system resilience through proactive reliability work.
- Find and mitigate single points of failure across distributed systems.
Production Operations & Incident Response
- Take part in and improve on-call rotations and incident response.
- Lead incident triage, mitigation, and resolution in real time.
- Run blameless postmortems and follow through on action items.
- Build tooling and automation to cut MTTR (Mean Time to Recovery).
Observability & System Insight
- Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
- Improve signal quality to cut noise and alert fatigue.
- Build dashboards and alerts that reflect real system health and user impact.
- Use observability data to drive performance and reliability improvements.
Performance & Scalability
- Analyze system performance under load and find bottlenecks.
- Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
- Partner with engineering teams to improve system efficiency and scaling behavior.
Automation & Reliability Tooling
- Build automation that eliminates repetitive operational work.
- Improve deployment safety through reliability checks and safeguards.
- Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
- Build tools for incident response, debugging, and capacity planning.
Security, Compliance & Operational Maturity
- Partner with security and compliance to keep systems meeting operational standards.
- Support audit readiness and reliability-related compliance requirements (Vanta).
- Integrate monitoring and alerting into security and SIEM workflows.
- Help mature operational practices across engineering.
You'll know it's working when SLOs and error budgets are clear and enforced, incidents are rare and shrinking over time, engineers trust their signals about system health, alerts are actionable instead of noisy, systems scale predictably under load, postmortems drive real improvement, and reliability is a shared responsibility, not a reactive function.
You May Be a Fit If
- You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
- You've operated and debugged distributed systems in production.
- You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
- You've defined and worked with SLOs, SLIs, and error budgets.
- You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
- You can write code or scripts (Python, Bash, etc.) for automation and tooling.
- You think in systems and reason clearly about failure modes.
- Bonus points for experience in high-growth or high-scale environments, background in regulated industries like healthcare or fintech, experience with ClickHouse or analytical systems at scale, familiarity with chaos engineering or load testing, and exposure to ML infrastructure or data platforms.
Why You'll Love Working Here
- Mission-Driven, World-Class Team — Join an exceptional group of professionals aligned around a meaningful mission and committed to making an impact
- Opportunities for Growth — Strengthen your expertise through collaboration with experienced, high-performing leaders across the organization
- Flexible Hybrid Work Environment — We're remote-first, with meaningful office presence in San Francisco and New York. R&D roles follow a hybrid model, with two days per week in our San Francisco office
Benefits & Perks
- Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family
- 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
- Equity — Every full-time employee shares in our success
- Unlimited PTO — Take the time you need, when you need it
- Daily Lunch Stipend — $100/week to cover your midday meals
- Wellness Stipend — $100/month to support your health and well-being
- Commuter Benefits — $100/month for SF and NYC-based employees
- Parental Leave — Paid leave to support growing families
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating...SeniorFull time- ...love to meet you. Our Enterprise Information Technology (EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this role, you will move beyond traditional infrastructure...SeniorPermanent employmentFull timeH1bLocal areaRemote workShift work
$182.8k - $247.3k
...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement$158.5k - $172k
...the exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate,... ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire...SeniorFull timeTemporary workWork at officeFlexible hours3 days per week$141k - $216.6k
...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical...SeniorWork experience placementWork at office$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...SeniorFull timeTemporary workLocal areaRemote workWorldwide$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....SeniorLocal area
$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience Technologies...SeniorFull timeLocal areaRemote workRelocation packageFlexible hoursShift work$500 per month
...accounts. Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to... ...significant impact, we encourage you to apply. Your Role: As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform...SeniorHome office$115k - $160k
...issues. We need strong systems engineering expertise, with deep... ...effective fixes that preserve reliability and performance. We require... ...management skills for engaging with senior business and technology... ...innovative solutions. This Senior Site Reliability Engineer - AVP -...SeniorFull timeWork at office$191k - $226k
....S. — and using AI to scale that impact further and faster than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads...SeniorRemote workWork visaFlexible hours- ...Job Description A major financial services company in NYC is growing its team rapidly, and they are looking for a Senior DevOps Engineer / Site Reliability Engineer who can join. If you’re passionate about high-availability, reliability, automation, we’d be excited...Senior
- ...We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs. The InfraSec team collaborates...SeniorFull timeRemote workWorldwide
- ...Title: Sr. Site Reliability Engineer (SRE) Location: New York City, NY - LOCALS ONLY Work Arrangement: Hybrid, 3 days Duration: 6-... ...Experience Range: 10-15 years Our client is seeking a Senior Site Reliability Engineer (SRE) with 10-15 years of experience...SeniorContract workLocal area
$104k - $178k
...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The team is responsible for building and maintaining the reliability, scalability, and performance of DV's digital media measurement...Senior$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ...while embedding directly with product engineering teams and their leadership to drive... ...eliminated manual deployment processes.At the Senior Engineer level, your influence is...SeniorFull timeWork experience placementWork at officeLocal area2 days per week3 days per week- ...Principal Site Reliability Engineer Location: New York, NY (Onsite) Job type: Contract Job Description: Job Requirements Must Have: - Site Reliability Engineering and system reliability optimization - GitLab CI/CD and HashiCorp Vault secrets management -...Contract work
$400k
...in financial markets, the organization combines innovation, engineering excellence, and data-driven insights to support complex trading operations worldwide. This opportunity is for a Senior Site Reliability Engineer to join a high-performance infrastructure...SeniorPermanent employmentWorldwide$168k - $200k
...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,...Senior$207k - $300k
...providing feedback to ensure best practices in reliability, security, and efficiency.Triage and... ...development initiatives. Mentor other engineers and contribute to the engineering... ...principles to cloud environments; and Managing senior stakeholders, external partners, and...Full timeWork at office- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad...Shift work
$45 - $85 per hour
DescriptionThe Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.We're looking for engineers to be part of an empowered, self-organizing group, with the opportunity to use modern languages and tools and to operate software...Contract workTemporary work$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure...Work at officeLocal areaImmediate start- ...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the... ...crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background....SeniorFull timeRemote workWorldwide
$150k - $190k
DescriptionKforce has a client that is seeking a Senior Principal Software Engineer (Delivery & Architecture) in New York, NY.Overview:We are seeking... ...in ambiguous environments and prioritizes high-quality, reliable delivery.Key Responsibilities:* Oversee the end-to-end...Senior$150k - $220k
...teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative... ...can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping...Local area$182k - $250.8k
...the backbone of our platform's reliability and operational excellence.... ...a forward-thinking group of engineers and leaders who believe that... ...users worldwide. As a Manager, Site Reliability Engineer, you'll... ...learningRepresent reliability as a senior technical leader in...Permanent employmentLocal areaRemote workWorldwideFlexible hoursWeekend workWeekday work$190k - $260k
...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving...Full timeWork experience placementWork at officeLocal areaRemote workHome office$150k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people... ...markets.Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience,...Full timeTemporary workPart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- site reliability engineer sre New York, NY
- senior technical analyst New York, NY
- senior associate attorney New York, NY
- senior developer New York, NY
- sr industrial security specialist New York, NY
- senior aws cloud engineer New York, NY
- senior manager business development New York, NY
- remote senior salesforce administrator New York, NY



