Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Plenful

Senior Site Reliability Engineer (SRE)

Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow.

This role is centered on operating real systems at scale — not just building infrastructure, but understanding deeply how it behaves under load, fails in production, and recovers. You'll define reliability standards, own production health, and build the feedback loops that make our systems more resilient over time.

You'll work closely with backend, data, and ML engineers to keep the platform highly available, measurable, and continuously improving — from incident response and performance debugging to SLO design and system-level optimization. This role is hybrid.

What You'll Do

Reliability Engineering & System Ownership

  • Define and implement SLIs, SLOs, and error budgets across core services.
  • Own production system health: uptime, latency, and availability targets.
  • Improve system resilience through proactive reliability work.
  • Find and mitigate single points of failure across distributed systems.

Production Operations & Incident Response

  • Take part in and improve on-call rotations and incident response.
  • Lead incident triage, mitigation, and resolution in real time.
  • Run blameless postmortems and follow through on action items.
  • Build tooling and automation to cut MTTR (Mean Time to Recovery).

Observability & System Insight

  • Design and evolve observability across metrics, logs, and distributed tracing (OpenTelemetry), using tools like Datadog, CloudWatch, Grafana, and Sentry.
  • Improve signal quality to cut noise and alert fatigue.
  • Build dashboards and alerts that reflect real system health and user impact.
  • Use observability data to drive performance and reliability improvements.

Performance & Scalability

  • Analyze system performance under load and find bottlenecks.
  • Optimize latency, throughput, and resource use across serverless (AWS Lambda), containerized services (ECS), and data systems (Aurora Postgres, ClickHouse).
  • Partner with engineering teams to improve system efficiency and scaling behavior.

Automation & Reliability Tooling

  • Build automation that eliminates repetitive operational work.
  • Improve deployment safety through reliability checks and safeguards.
  • Contribute to CI/CD pipelines (GitHub Actions) with a focus on stability.
  • Build tools for incident response, debugging, and capacity planning.

Security, Compliance & Operational Maturity

  • Partner with security and compliance to keep systems meeting operational standards.
  • Support audit readiness and reliability-related compliance requirements (Vanta).
  • Integrate monitoring and alerting into security and SIEM workflows.
  • Help mature operational practices across engineering.

You'll know it's working when SLOs and error budgets are clear and enforced, incidents are rare and shrinking over time, engineers trust their signals about system health, alerts are actionable instead of noisy, systems scale predictably under load, postmortems drive real improvement, and reliability is a shared responsibility, not a reactive function.

You May Be a Fit If
  • You've spent 5+ years in Site Reliability Engineering, SRE-adjacent roles, or production infrastructure.
  • You've operated and debugged distributed systems in production.
  • You have hands-on experience with observability tooling (Datadog, Grafana, OpenTelemetry, or similar), incident response and on-call practices, and performance and reliability debugging.
  • You've defined and worked with SLOs, SLIs, and error budgets.
  • You're familiar with AWS environments, serverless and container-based architectures, and Postgres or similar relational databases.
  • You can write code or scripts (Python, Bash, etc.) for automation and tooling.
  • You think in systems and reason clearly about failure modes.
  • Bonus points for experience in high-growth or high-scale environments, background in regulated industries like healthcare or fintech, experience with ClickHouse or analytical systems at scale, familiarity with chaos engineering or load testing, and exposure to ML infrastructure or data platforms.
Why You'll Love Working Here
  • Mission-Driven, World-Class Team — Join an exceptional group of professionals aligned around a meaningful mission and committed to making an impact
  • Opportunities for Growth — Strengthen your expertise through collaboration with experienced, high-performing leaders across the organization
  • Flexible Hybrid Work Environment — We're remote-first, with meaningful office presence in San Francisco and New York. R&D roles follow a hybrid model, with two days per week in our San Francisco office
Benefits & Perks
  • Healthcare Coverage — Full medical, dental, and vision insurance for you and participation for your family
  • 401(k) with Company Match — Plenful matches 50% of your first 3% contributed
  • Equity — Every full-time employee shares in our success
  • Unlimited PTO — Take the time you need, when you need it
  • Daily Lunch Stipend — $100/week to cover your midday meals
  • Wellness Stipend — $100/month to support your health and well-being
  • Commuter Benefits — $100/month for SF and NYC-based employees
  • Parental Leave — Paid leave to support growing families
Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in New York, NY vacancy
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating... 
    Senior
    Full time

    Ridgeline

    New York, NY
    1 day ago
  •  ...love to meet you. Our Enterprise Information Technology (EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this role, you will move beyond traditional infrastructure... 
    Senior
    Permanent employment
    Full time
    H1b
    Local area
    Remote work
    Shift work

    Jack Henry & Associates

    New York, NY
    1 day ago
  • $182.8k - $247.3k

     ...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed... 
    Senior
    Work experience placement

    Duolingo

    New York, NY
    6 days ago
  • $158.5k - $172k

     ...the exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate,...  ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    4 days ago
  • $141k - $216.6k

     ...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical... 
    Senior
    Work experience placement
    Work at office

    Axon

    New York, NY
    2 days ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Senior
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    6 days ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    3 days ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Senior
    Local area

    E-Solutions

    New York, NY
    4 days ago
  • $189k - $283.6k

     ...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics...  ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience Technologies... 
    Senior
    Full time
    Local area
    Remote work
    Relocation package
    Flexible hours
    Shift work

    Block USA

    New York, NY
    3 days ago
  • $500 per month

     ...accounts. Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to...  ...significant impact, we encourage you to apply. Your Role: As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform... 
    Senior
    Home office

    Alpaca

    New York, NY
    9 days ago
  • $115k - $160k

     ...issues. We need strong systems engineering expertise, with deep...  ...effective fixes that preserve reliability and performance. We require...  ...management skills for engaging with senior business and technology...  ...innovative solutions. This Senior Site Reliability Engineer - AVP -... 
    Senior
    Full time
    Work at office

    Barclays

    New York, NY
    3 days ago
  • $191k - $226k

     ....S. — and using AI to scale that impact further and faster than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads... 
    Senior
    Remote work
    Work visa
    Flexible hours

    Garner Health

    New York, NY
    8 days ago
  •  ...Job Description A major financial services company in NYC is growing its team rapidly, and they are looking for a Senior DevOps Engineer / Site Reliability Engineer who can join. If you’re passionate about high-availability, reliability, automation, we’d be excited... 
    Senior

    The Greene Group

    New York, NY
    15 days ago
  •  ...We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs. The InfraSec team collaborates... 
    Senior
    Full time
    Remote work
    Worldwide

    Mongodb

    New York, NY
    2 days ago
  •  ...Title: Sr. Site Reliability Engineer (SRE) Location: New York City, NY - LOCALS ONLY Work Arrangement: Hybrid, 3 days Duration: 6-...  ...Experience Range: 10-15 years Our client is seeking a Senior Site Reliability Engineer (SRE) with 10-15 years of experience... 
    Senior
    Contract work
    Local area

    RIT Solutions, Inc.

    New York, NY
    3 days ago
  • $104k - $178k

     ...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The team is responsible for building and maintaining the reliability, scalability, and performance of DV's digital media measurement... 
    Senior

    DoubleVerify

    New York, NY
    3 days ago
  • $140k - $215k

     ...intersection of our Core Platform and Embedded Reliability charters: building the foundational...  ...while embedding directly with product engineering teams and their leadership to drive...  ...eliminated manual deployment processes.At the Senior Engineer level, your influence is... 
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    New York, NY
    4 days ago
  •  ...Principal Site Reliability Engineer Location: New York, NY (Onsite) Job type: Contract Job Description: Job Requirements Must Have: - Site Reliability Engineering and system reliability optimization - GitLab CI/CD and HashiCorp Vault secrets management -... 
    Contract work

    Argyle Infotech

    New York, NY
    4 days ago
  • $400k

     ...in financial markets, the organization combines innovation, engineering excellence, and data-driven insights to support complex trading operations worldwide. This opportunity is for a Senior Site Reliability Engineer to join a high-performance infrastructure... 
    Senior
    Permanent employment
    Worldwide
    New York, NY
    more than 2 months ago
  • $168k - $200k

     ...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,... 
    Senior

    Datavant

    New York, NY
    23 days ago
  • $207k - $300k

     ...providing feedback to ensure best practices in reliability, security, and efficiency.Triage and...  ...development initiatives. Mentor other engineers and contribute to the engineering...  ...principles to cloud environments; and Managing senior stakeholders, external partners, and... 
    Full time
    Work at office

    Google

    New York, NY
    5 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad... 
    Shift work

    JP Morgan Chase

    New York, NY
    5 days ago
  • $45 - $85 per hour

    DescriptionThe Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.We're looking for engineers to be part of an empowered, self-organizing group, with the opportunity to use modern languages and tools and to operate software... 
    Contract work
    Temporary work

    TEKsystems

    New York, NY
    4 days ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    4 days ago
  •  ...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the...  ...crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background.... 
    Senior
    Full time
    Remote work
    Worldwide

    Mongodb

    New York, NY
    a month ago
  • $150k - $190k

    DescriptionKforce has a client that is seeking a Senior Principal Software Engineer (Delivery & Architecture) in New York, NY.Overview:We are seeking...  ...in ambiguous environments and prioritizes high-quality, reliable delivery.Key Responsibilities:* Oversee the end-to-end... 
    Senior

    KForce

    New York, NY
    3 days ago
  • $150k - $220k

     ...teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative...  ...can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping... 
    Local area

    Forge Global

    New York, NY
    3 days ago
  • $182k - $250.8k

     ...the backbone of our platform's reliability and operational excellence....  ...a forward-thinking group of engineers and leaders who believe that...  ...users worldwide. As a Manager, Site Reliability Engineer, you'll...  ...learningRepresent reliability as a senior technical leader in... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    New York, NY
    4 days ago
  • $190k - $260k

     ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    4 days ago
  • $150k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people...  ...markets.Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience,... 
    Full time
    Temporary work
    Part time

    Goldman Sachs

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!