Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Cellebrite

Description

About Cellebrite:

Cellebrite's (Nasdaq: CLBT) mission is to enable its global customers to protect and save lives by enhancing digital investigations and intelligence gathering to accelerate justice in communities around the world. Cellebrite's AI-powered Digital Investigation Platform enables customers to lawfully access, collect, analyze and share digital evidence in legally sanctioned investigations while preserving data privacy. Thousands of public safety organizations, intelligence agencies and businesses rely on Cellebrite's digital forensic and investigative solutions-available via cloud, on-premises and hybrid deployments-to close cases faster and safeguard communities.

To learn more, visit us at and find us on social media @Cellebrite.


What is your mission?


As a Site Reliability Engineer, you'll be a key technical owner of production health for Cellebrite's cloud platform the platform investigators and public safety agencies depend on every day. You'll work closely with our TCS support team and R&D to triage, investigate, and resolve production issues quickly and thoroughly, while helping modernize how we monitor the platform, roll out changes safely, and respond when things break. The focus of this role is keeping production healthy and continuously raising the bar on how we operate it - not building new infrastructure from scratch. Your work directly supports Cellebrite's mission to protect lives, accelerate justice, and preserve data privacy.


Responsibilities:

Incident Response & Production Health

  • Own the full incident lifecycle detect, triage by severity/impact, investigate, and drive to resolution engaging TCS, R&D, and DevOps as needed.
  • Act as a technical responder during major incidents and grow into an incident-commander role coordinating the response across teams in real time.
  • Lead blameless post-incident reviews and make sure corrective actions actually get implemented, not just documented.
  • Track recurring issues and support-ticket trends; distinguish patterns that need a permanent fix from one-off noise, and route the former to R&D.
AI-Driven Monitoring & Observability
  • Evolve production monitoring toward AI-assisted operations anomaly detection, cross-signal correlation across logs/metrics/traces, and LLM-based triage assistants that cut time-to-diagnosis.
  • Continuously tune dashboards, alert thresholds, and routing so real issues surface fast and noise doesn't drown them out.
  • Evaluate and pilot AI/ML-based observability tooling, and champion adoption of tools like Copilot or log-analysis assistants across the team.
Rolling Upgrades & Change Management
  • Plan and execute rolling upgrades, version updates, and patches to production services and infrastructure with zero or minimal downtime.
  • Partner with R&D and DevOps to define safe rollout and rollback strategies for deployments, and validate system health post-upgrade.
  • Participate in an on-call rotation to help maintain production uptime SLOs, with upgrade windows and rollback readiness built into the runbooks you own.
Cross-Team Enablement
  • Partner with TCS on production tickets, providing deeper technical investigation when issues exceed their level; partner with R&D on code-level fixes, deployments, and architectural input.
  • Own and continuously improve runbooks and the known-issues knowledge base so TCS resolves more independently over time.
  • Maintain clear, current documentation of production architecture, known issues, and resolution paths for both TCS and R&D.
Requirements

Requirements:
  • 3-5 years of experience in a production support, SRE, or operations role.
  • Solid working experience with AWS (troubleshooting and operating existing infrastructure; deep infra design experience is not required).
  • Experience with monitoring/observability tools (e.g., Datadog, CloudWatch, Grafana) able to read dashboards, tune alerts, and investigate from logs/metrics/traces.
  • Working knowledge of Linux system administration and basic networking troubleshooting.
  • Familiarity with containerized environments (Kubernetes) - enough to check pod health, logs, and restart/rollback safely.
  • Comfortable with scripting (Python or Bash) to automate repetitive troubleshooting or reporting tasks.
  • Experience driving rolling upgrades, patching, or zero-downtime deployment processes in a production environment.
  • Experience collaborating across support and engineering teams - comfortable working with an outsourced support team (e.g., TCS) on one side and R&D engineers on the other, translating between the two.
  • Strong communication skills in English, written and verbal - clear handoffs and documentation matter as much as fixing things.
  • Nice to have: exposure to Terraform or other IaC (read/troubleshoot, not necessarily author), basic CI/CD familiarity, and genuine interest in applying AI/ML to monitoring, alerting, and incident triage - not just using AI coding assistants.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in New York, NY vacancy
  •  ...investing smarter, simpler, and more accessible for everyone. As part of our global expansion, we’re looking for a hands-on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our next-generation financial platforms. This is a high-impact... 
    Suggested
    Full time
    Remote work

    Longbridge

    New York, NY
    3 days ago
  •  ...We are seeking a Kubernetes Engineer to join our team. In this role, you will be instrumental in shaping and executing our containerization strategy, optimizing resource utilization, and ensuring robust disaster recovery and business continuity. The ideal candidate has... 
    Suggested
    Work at office

    Edge Search formerly Alpha Search Advisors

    New York, NY
    3 days ago
  •  ...Senior Site Reliability Engineer We're partnering with an elite Prime Brokerage firm looking for a high-calibre Site Reliability Engineer to help build and scale the infrastructure underpinning a global trading platform. This is a software-focused SRE role for engineers... 
    Suggested
    Full time

    Radley James

    New York, NY
    2 days ago
  •  ...A leading High-Frequency Trading firm is seeking an experienced Senior Site Reliability Engineer. This is a critical role where reliability, performance, automation and operational excellence are essential. You will work closely with software engineers, infrastructure... 
    Suggested
    Full time

    Radley James

    New York, NY
    2 days ago
  •  ...About the job A leading quantitative trading firm is seeking a Senior Site Reliability Engineer to build and evolve the reliability, observability, and automation capabilities powering a highly performance-sensitive trading environment. Working at the intersection... 
    Suggested
    Full time

    Acquire Me

    New York, NY
    20 hours ago
  • $120k - $150k

     ...allows each person to achieve personal success and add value to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About the RoleJoin our Platform Engineering team as a Site Reliability... 
    Full time

    Piper Sandler Companies

    New York, NY
    4 days ago
  •  ...The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also... 
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    5 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    3 days ago
  •  ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 

    re-tool®

    New York, NY
    3 days ago
  •  ...Job Description Chariot’s engineering hire will be responsible for taking the Chariot platform to the next level. You will lead the evolution of our banking product, DAF product and technology strategy, along with building a growing team of software engineers. You will... 
    Work experience placement
    Work at office
    Work from home
    Monday to Friday
    Monday to Thursday

    Chariot

    New York, NY
    3 days ago
  • $150k - $200k

     ...Defence and Government. They’re continuing their expansion of their New York (and Washington DC) engineering teams and looking for an Infrastructure Engineer / Site Reliability Engineer / Forward Deployed Infrastructure Engineer with a strong software engineering... 

    JLA Resourcing Ltd

    New York, NY
    3 days ago
  • $100k - $250k

     ...systems) can be yours. What you’ll do Improve observability, reliability and availability by defining and measuring key metrics....  ...function improvements. Educate, mentor and hold accountable the engineering team to improve the reliability of our systems and make... 
    Local area

    Kalshi

    New York, NY
    3 days ago
  • $182.8k - $247.3k

     ...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo's sophisticated distributed systems... 
    Work experience placement

    Drive Capital

    New York, NY
    3 days ago
  • $171.6k - $223k

     ...development and learning. It allows us to scale easily, enabling our engineers to enhance attention on new features and capabilities. A key...  ...of members all over the world. Peloton is looking for a Site Reliability Engineer to create the tooling and services which simplify... 
    Temporary work
    Local area

    Peloton

    New York, NY
    4 days ago
  •  ...Improve the reliability of mission critical solutions, applications, and platforms Software development for enterprises Continuous...  ...Shell Languages Powershell and Bash Windows and Linux Years of Experience: 5 Years of Software Engineering #J-18808-Ljbffr
    Work experience placement

    InterEx Group

    New York, NY
    2 days ago
  • $104k - $178k

    ## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR00000779# ****Who We Are****DV is the leader in digital performance solutions, helping our advertiser and agency partners Verify the quality of their digital campaigns, Optimise... 
    Full time

    DoubleVerify

    New York, NY
    1 day ago
  •  ...firm that has been operating at scale since 2014, with millions of global users and a reputation for rigorous engineering. The Role As Senior Site Reliability Engineer , you will own the infrastructure foundation that the entire engineering organization depends on.... 

    Harrison Clarke

    New York, NY
    3 days ago
  •  ...DEPARTMENT: Product Engineering / Operational Readiness REPORTING TO: Senior Manager, System...  ...Engineering team is responsible for the reliability, monitoring, automation, and operational...  ...services. Role Overview: The Site Reliability Engineer II (SRE) is responsible... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    IPC Systems, Inc.

    New York, NY
    3 days ago
  •  ...Ireland Full time Are you passionate about building reliable, scalable systems that power critical business solutions?...  ...can learn more about LexisNexis Risk at the Role As a Site Reliability Engineer (SRE), you will bridge software development and IT operations... 
    Full time
    Work from home

    RELX

    New York, NY
    3 days ago
  •  ...on one unified cloud. One cloud for compute, inference, and agents. Role Overview We are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and... 

    GMI Cloud

    New York, NY
    9 hours ago
  •  ...Traders, Tower Research, PDT Partners, SIG, and more. We're looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability of our platform, and help set the standard for how... 
    Full time
    Local area
    Remote work

    Databento

    New York, NY
    2 days ago
  • $120k - $160k

     ...and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more. As a Site Reliability Engineer (SRE), you will work at the intersection of production operations and software development as you improve, manage, and... 
    Work at office
    Local area

    The Voleon Group

    New York, NY
    3 days ago
  •  ...Responsibilities Support and enhance the reliability, availability, and performance of...  ...service improvements. Work closely with engineering teams to implement, test, and deploy...  ...production environments, platform engineering, site reliability, DevOps, or infrastructure... 

    Selby Jennings

    New York, NY
    3 days ago
  • $100k - $135k

     ...Model-Based Manufacturing, where context-aware production planning informs design in real-time—eliminating the disconnect between engineering and manufacturing. Our platform automates what can be automated and captures tribal knowledge where automation falls short,... 
    Work at office
    Flexible hours

    Dirac, Inc.

    New York, NY
    3 days ago
  • $86.13k - $127.19k

     ...and build a more sustainable, more inclusive world.Job Role:Job Role: Site Reliability EngineerLocation: Boston, MADuration: FulltimeSummary:We are seeking a highly motivated Site Reliability Engineer to help build and operate reliable, scalable, and secure services... 
    Full time
    Local area

    Capgemini

    New York, NY
    2 days ago
  •  ...CTG is seeking to fill a Site Reliability Engineer opening for our client in New York, NY. Location: New York, NY Duration: 12 months Duties Operates company's internal data communications systems. Plans, designs and implements local and wide-area network... 
    Local area

    Computer Task

    New York, NY
    1 day ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Local area

    E-Solutions

    New York, NY
    3 days ago
  • $145k - $190k

     ...slowly. The product is live. The market is growing. The category is still ours to define. The Role. We're hiring a Site Reliability Engineer to build the infrastructure, systems, and operational practices that keep Onyx fast and reliable as we scale. Markets move... 
    Permanent employment
    Work at office
    Immediate start

    Onyxodds

    New York, NY
    1 day ago
  • $250k - $300k

     ...skills and experience — talk with your recruiter to learn more. Base pay range $250,000.00/yr - $300,000.00/yr Senior Site Reliability Engineer – Trading Systems This isn’t a support role. It’s an engineering position where uptime, performance, and automation... 
    Full time
    Work at office
    Home office

    Quantitative Systems

    New York, NY
    3 days ago
  • $115k - $125k

     ...headquartered in New York, with offices in Chicago, London, Singapore, Hong Kong and Tokyo. Purpose of the role Pico Site Reliability Engineering is a customer-facing group engaged with our customers, development and production teams to ensure customer success with... 
    Contract work
    For contractors
    For subcontractor
    Work at office
    Work from home
    Monday to Friday
    Flexible hours
    Shift work
    Weekend work
    Afternoon shift
    Early shift

    Pico

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!