Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, AI Observability

Appnovation

About usAppnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth. Bold ambition. Practical action. Endless possibilities.We’re looking for a Site Reliability Engineer to keep a shared observability platform for LLM-based applications running for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, with ClickHouse as the analytical store, PostgreSQL for metadata and Redis as the ingestion queue, all delivered through Argo CD.Two things need building rather than maintaining. The platform has no monitoring, alerting or defined service levels today, and you will own putting them in place. Infrastructure is defined as code throughout, in Kubernetes manifests, Helm values and Argo CD applications.Alongside the platform itself, you will own the runbooks that make an incident survivable by someone other than the author, and the onboarding and support path for the internal teams that depend on the platform.ROLE RESPONSIBILITIESPlatform Operations: Diagnose and resolve failures across ClickHouse, PostgreSQL, Redis, the ingestion workers and the Kubernetes layer beneath them, including ingest backpressure from queue depth and worker drain behaviour.ClickHouse Operations: Own ClickHouse under the operator model, including Keeper quorum, replication, shard and replica topology, and S3 storage tiering.Provable Backup and Restore: Rehearse the restore, time it, document it and test it against its failure modes, rather than assuming a successful backup job means a recoverable system.Safe Upgrades: Plan and rehearse upgrades in a lower environment, with a rollback plan that works even if a migration is only partly complete.Monitoring and Service Levels: Build monitoring, alerting and service levels from scratch, so problems are found here before a user reports them.Infrastructure as Code: Maintain Kubernetes manifests, Helm values and Argo CD applications so every change goes through the delivery pipeline.Runbooks and SOPs: Write and maintain runbooks and SOPs that let a colleague resolve an incident without the author present.Onboarding and Support: Run the onboarding and support path for internal teams that depend on the platform, and triage what they bring.Automation: Turn recurring operational work into automation.Upgrade Partnership: Work with the platform engineer who owns what the platform offers. They decide what to adopt and how it is configured; you own the migration and its rollback.QUALIFICATIONSHands-on experience with Kubernetes on AWS (managed EKS), with routine work done through Helm values and Argo CD applications.Hands-on experience running ClickHouse in production, including replication and Keeper quorum, shard and replica topology, and backup and restore, ideally run through an operator.Experience building monitoring and alerting from scratch, including service levels that reflect what users actually experience rather than what is easy to measure.Experience upgrading self-hosted software safely, including schema migrations, rehearsal in a lower environment and a rollback plan for a partly completed migration.PostgreSQL and Redis operations deep enough to debug metadata-store and queue problems, including backpressure and worker drain.Strong operational writing: runbooks, SOPs and post-incident reviews that a colleague can follow unaided during an incident.PREFERRED QUALIFICATIONSObservability engineering, including OpenTelemetry Collector pipelines, alerting design and Grafana dashboards.OIDC or enterprise SSO integration with a corporate identity provider.GitHub Actions for plan and apply pipelines with approval gates.Experience running LLM observability tools such as Langfuse, LangSmith or Arize Phoenix.Experience in pharma, life sciences or another regulated industry.WHO YOU AREYou don’t trust a backup until you’ve restored from itYou want to find problems before users doYou write runbooks for the person on call at 3am, not for yourselfYou stay calm in incidents and focus on fixing the process afterwardYou automate anything you have to do twiceYou work well inside a client team and build trust quicklyYou have prior experience in consultingPrior experience and connections in the Life Sciences industry is preferredThank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital. Accommodations are available upon request throughout the recruitment process.

Vacancy posted 1 hour ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, AI Observability in New York, NY vacancy
  •  ...GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help...  ...particularly agentic development and AI-assisted engineering, and identify opportunities...  ..., Java, Go or RustExperience with observability and monitoring platforms such as... 
    Suggested
    Full time
    Work experience placement
    Remote work

    Shutterstock

    New York, NY
    3 days ago
  • $200k - $250k

     ...River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our...  ...systems - from directory services to AI tools, both on-prem and in the cloud...  ...skill set in Linux, Kubernetes, and observability; proficiency in Python; and solid experience... 
    Suggested
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    2 days ago
  • $120k - $150k

     ...communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team...  ..., databases, storage, and growing AI workloads — applying DevSecOps and...  ...respond to incidents using enterprise observability tooling, contributing to alerting,... 
    Suggested
    Full time

    Piper Sandler Companies

    New York, NY
    2 days ago
  •  ...complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the...  ...pipelinesUses enterprise-authorized AI capabilities within the work environment...  ...issues emerge.Enhance production observability and reliability by improving instrumentation... 
    Suggested
    Shift work

    JP Morgan Chase

    New York, NY
    2 days ago
  •  ...realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the...  ...practices, expanding SRE adoption (observability, monitoring, automation, operational...  ...).Uses enterprise-authorized AI capabilities within the work environment... 
    Suggested

    JP Morgan Chase

    New York, NY
    4 days ago
  • $190k - $260k

     ...security-first enterprise AI company. We build...  ...team of researchers, engineers, designers, and more,...  ...performance, scalable and reliable machine learning...  ...We are looking for a Site Reliability Engineer to...  ....Automate environment observability and resilience. Enable... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    1 day ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential...  ...and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in...  ...service communication, security, and observability within the Kubernetes clusters. Enable... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    2 days ago
  • $130k - $190k

     ...Insurance and Payments verticals, supplemented with Data & AI, Cybersecurity, Risk & Compliance, Change Management and Digital...  ...the Role: Summary Responsible for the operational reliability, observability, and stability of the Strategic Full Revaluation Capability... 
    Temporary work
    Remote work
    Worldwide

    BIP US

    New York, NY
    2 days ago
  • $150k - $250k

     ...We DoAt Goldman Sachs, our Engineers don't just make things - we...  ...Banking & Markets business, the Site Reliability Engineering (SRE) team...  ..., cloud infrastructure, and AI-driven operations. Using Goldman...  ...security and comprehensive observability (metrics, distributed tracing... 
    Full time
    Temporary work
    Part time

    Goldman Sachs

    New York, NY
    10 hours ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential...  ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and...  ...world class, comprehensive, scalable Observability Platform that enables our SRE teams and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    10 hours ago
  • $182k - $250.8k

     ...Every Identity, from AI to HumanIdentity is the...  ...backbone of our platform's reliability and operational...  ...forward-thinking group of engineers and leaders who...  ...worldwide. As a Manager, Site Reliability Engineer,...  ...standards that embed observability, resilience, and software... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    New York, NY
    2 days ago
  • $104k - $178k

    ## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR000...  ...technical contributor driving automation, observability, and operational excellence across...  ...measurement platforms.* Leverage AI-assisted development tools to accelerate... 
    Full time

    DoubleVerify

    New York, NY
    4 days ago
  • $86.13k - $127.19k

     ...inclusive world.Job Role:Job Role: Site Reliability EngineerLocation: Boston,...  ...motivated Site Reliability Engineer to help build and operate...  ...excellence through automation, observability, and engineering best...  ...market leading capabilities in AI, generative AI, cloud and data... 
    Full time
    Local area

    Capgemini

    New York, NY
    10 hours ago
  •  ...DEPARTMENT: Product Engineering / Operational Readiness...  ...is responsible for the reliability, monitoring, automation...  ...Role Overview: The Site Reliability Engineer...  ...enterprise monitoring, observability, event management, and...  ...with responsible use of AI-assisted engineering... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    IPC Systems, Inc.

    New York, NY
    1 day ago
  •  ...empowering our clients to succeed in an evolving digital landscape. Role Overview We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. In this role, you will... 

    Ontrac Solutions Inc

    New York, NY
    4 days ago
  • $234k - $300k

    The ML Observability team builds cutting-edge tools to monitor, explain, and improve AI systems in production, particularly those leveraging...  ...with confidence.As a Staff Engineer, you’ll lead the development...  ..., understandable, and reliable in the real world.At Datadog... 
    Work at office

    Datadog

    New York, NY
    1 day ago
  • $244k - $305k

     ...is seeking a Staff Software Engineer to help shape the future of our...  ...) Logs offering by unifying observability pipelines with log management...  ...observability, diagnostics, reliability, and secure operation across...  ...and security platform for the AI era, providing businesses with... 
    Work at office

    Datadog

    New York, NY
    3 days ago
  • $200k - $240k

     ...Senior Site Reliability Engineer Inside health systems, where every second can matter, operations...  ...We combine proprietary hardware, AI-powered intelligence, and deep integrations...  ...when things break Build out observability that's genuinely tuned — SLIs, SLOs,... 
    Work at office
    3 days per week

    Kontakt.io

    New York, NY
    3 days ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer...  ...systems, processes, automations, and observability tooling that keep our platform... 
    Flexible hours

    Baseten

    New York, NY
    10 hours ago
  • $189k - $283.6k

     ...proactively and reactively improve the reliability of Block's platform and...  ...leverage and continuously improve AI-driven tooling and automation to enhance observability, accelerate incident detection...  ...to perform and grow as an engineer ~5+ years of software development... 
    Full time
    Local area
    Remote work
    Relocation package
    Flexible hours
    Shift work

    Block USA

    New York, NY
    10 hours ago
  •  ...Mistral Mistral provides full-stack AI solutions: from frontier models to...  ...We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability...  ...working against reliability KPIs (observability, alerting, SLAs) • Hands-on... 
    Relocation package

    Mistral Ai

    New York, NY
    1 day ago
  • $111k - $218k

     ...Site Reliability Engineer The Site Reliability Engineering team designs and builds the global infrastructure...  ...infrastructure. Experience with observability of large scale distributed systems....  ...redefined the data platform for the AI era, enabling builders to create,... 
    Local area
    Worldwide

    Realm, Inc

    New York, NY
    1 day ago
  • $175k - $225k

     ...Site Reliability Engineer New York, New York, United States The Role We're looking for a Site Reliability Engineer to join our AI Technology team and play a critical role in ensuring the...  ...code, automation workflows, observability, metrics, and support users as... 

    Schonfeld

    New York, NY
    1 day ago
  • $250k - $300k

     ...Get AI-powered advice on this job and more exclusive features. This range...  ...000.00/yr - $300,000.00/yr Senior Site Reliability Engineer – Trading Systems This isn’t a support...  ...production workflows. Deploy real observability and alerting no noise, just signal.... 
    Full time
    Work at office
    Home office

    Quantitative Systems

    New York, NY
    1 day ago
  • $100k - $250k

     ...economics, financials, weather, tech, AI, culture and more. We believe prediction...  ...Roadmap As a member of Kalshi's engineering team, you'll help build the next-...  ...evolve. What You'll Do Improve observability, reliability, and service availability by defining... 
    Local area

    Kalshi Inc

    New York, NY
    10 hours ago
  • $135k - $160k

     ...About the Role We are looking for a Site Reliability Engineer to help us evolve and safeguard the...  ...is also the easiest. Automation & AI-Assisted Operations Providing robust...  ...procedures and repetitive tasks. Observability & Incident Management Driving the... 

    fabrichealth

    New York, NY
    1 day ago
  • $120k - $140k

     ...Site Reliability Engineer – AWS, Azure, IaC, Typescript, .NET – NY (office 3 days a week) - $120,00...  ...support knowledge sharing. Improve observability through monitoring, alerting and dashboards...  ...or fintech sector. Exposure to AI/ML platforms, prompt engineering or... 
    Full time
    Work at office
    Rotating shift
    3 days per week

    Stott and May

    New York, NY
    4 days ago
  • $140k - $215k

     ...with the world’s most advanced AI-native platform. We work on...  ...Core Platform and Embedded Reliability charters: building the...  ...embedding directly with product engineering teams and their leadership to...  ...libraries and frameworks, maturing observability tooling (tracing, profiling,... 
    Full time
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    New York, NY
    1 day ago
  • $220k - $260k

     ...Tabs is the leading AI-native revenue platform for modern finance and accounting...  ...The Role We’re looking for a Staff Site Reliability Engineer to lead the evolution of Tabs’...  ...and operate systems that are reliable, observable, and easy to develop on. You’ll own... 
    Full time
    Contract work
    Work at office

    TABS inc.

    New York, NY
    1 day ago
  • $240k - $300k

     ...time Location Type On-site Department Engineering & Product Engineering...  ...step of the way. Our AI-native workspace empowers...  ...role As a Staff Site Reliability Engineer you'll play a lead...  ...and services Own the observability, capacity planning, and monitoring... 
    Full time
    Work at office

    Menlo Ventures

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, AI Observability. Be the first to apply!