Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Triomics

Triomics Backend Engineer

Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained staff manually reading unstructured patient records - pathology reports, clinical notes, genomic panels - to power workflows like trial matching, registry curation, visit prep, and quality reporting. We replace that manual work with task-driven AI agents that sit inside the EMR and process records at scale, in real time.

Our platform is trusted by the 4 of the top 10 Best Hospitals for Cancer by U.S.News and several of the largest community practices. We have grown 10x in the last year and process millions of oncology medical documents monthly.

Our investors include Lightspeed, General Catalyst, Nexus Venture Partners and Y-Combinator.

Role

This role spans backend product engineering and infrastructure. You'll build backend services and application features, and also own the cloud infrastructure, deployments, and CI/CD that keeps them running in production. The platform processes millions of clinical documents monthly across multi-tenant deployments in customer as well as Triomics cloud environments, with GPU infrastructure serving AI extraction models. We need someone who can write application code in the morning and debug a Kubernetes deployment issue in the afternoon.

What Success Looks Like in the First 90 Days

Days 1-30: Map the entire infrastructure and find what's fragile.

Get access to every deployment - AWS, Azure, customer-hosted environments. Understand the full topology: how Kubernetes clusters are configured, how GPU nodes serve models, how document pipelines move data from EHR ingestion to extraction to structured output. Your first job is to understand what is already built, where the sharp edges are, and what breaks when load spikes or a deployment goes sideways. By end of month one, you should have a written map of every production environment, know which deployments are most fragile, and have identified the top 3 infrastructure risks.

Days 30-60: Own production stability and start shipping backend services.

Take ownership of at least one customer deployment end-to-end - monitoring, alerting, incident response. Set up observability that catches pipeline failures and data quality regressions before customers report them (today, customers often find issues first). Simultaneously, pick up a backend product feature - patient data processing, document pipeline improvement, or a platform feature the product team needs. Ship it. The goal is to make sure you can context-switch between infra firefighting and product engineering.

Days 60-90: Standardize deployments and Monitor Everything.

Document deployment runbooks, automate what's manual, and build CI/CD improvements that make releases safer and faster. You should have a clear plan for what the infrastructure needs to look like to support 2-3x the current customer count without adding headcount proportionally.

Responsibilities
  • Build and ship infrastructure services that power our product - document pipelines, application logic, and platform features
  • Own cloud infrastructure and deployment pipelines across both Triomics and customer environments (AWS, Azure)
  • Manage Kubernetes clusters, containerized services, CI/CD, and release processes including GPU node management for model serving
  • Build monitoring, alerting, and observability across production deployments - we process millions of documents and need to catch pipeline failures, data quality regressions, and infrastructure issues before customers do
  • Debug and resolve production issues end-to-end - from application-layer bugs to infrastructure failures
  • A significant portion of our engineering team is offshore and this role requires working with that team as well on architecture decisions, code reviews, and production stability
Requirements
  • 3+ years as a platform/infrastructure engineer at a startup or growth-stage company
  • Strong backend engineering: can design, build, and ship production services
  • Comfortable across the infrastructure stack: cloud (AWS or Azure), Kubernetes, Docker, CI/CD, networking, monitoring
  • Experience managing production deployments and debugging issues across application and infrastructure layers.
  • Can context-switch between writing product code and doing infra/ops work without treating either as out of scope of their job
Preferred
  • Experience with data-heavy applications - document processing pipelines, batch and real-time data workflows
  • Worked with ML/AI systems in production - model serving, GPU infrastructure, pipeline orchestration
  • Built infrastructure at an early-stage company where you were one of few engineers owning the full stack
  • Familiarity with building third party integrations in product is a plus
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in New York, NY vacancy
  • $123k - $165k

    Job Summary:Department/Group OverviewOur engineering fleet is a horizontal set of teams...  ...organization. Our specific team provides reliability engineering and operational support to backend...  ...products and brands.We are seeking a Site Reliability Engineer who will contribute... 
    Suggested

    Disney Interactive

    New York, NY
    13 hours ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    1 day ago
  • $165k - $241.4k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Suggested
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    4 days ago
  • $141k - $216.6k

     ...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical... 
    Suggested
    Work experience placement
    Work at office

    Axon

    New York, NY
    3 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Suggested
    Full time
    Work at office
    3 days per week

    GrubHub

    New York, NY
    4 days ago
  • $138.1k - $198.2k

     ...more intuitive with technology that simply works.  The SRE Engineering Enablement Team supports our CI Platforms, Developer Environments...  ...Our customers are all engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter of our engineering... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    New York, NY
    4 days ago
  • $110k - $120k

     ...largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading financial services and healthcare technology company based on... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    1 day ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    1 day ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad... 
    Shift work

    JP Morgan Chase

    New York, NY
    1 day ago
  • $120k - $200k

     ...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery...  ...practices (e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Overseas

    Comrise

    New York, NY
    2 days ago
  • $104.9k - $174.7k

     ...Management. You can learn more about LexisNexis Risk at the link below, About the Role: We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    LexisNexis Risk Solutions

    New York, NY
    13 hours ago
  •  ...human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You... 
    Full time
    Work at office

    Dune Security

    New York, NY
    3 days ago
  • $140k - $170k

     ...Senior Site Reliability Engineer New York About us @Symphony Secure. Connected. Intelligent. Symphony is an AI-powered communication and technology company fueled by interconnected platforms: messaging, voice, directory and analytics. Our end-to-end encrypted... 
    Local area

    Symphony Service Corp

    New York, NY
    1 day ago
  •  ...A financial technology company based in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability and health of their fintech products. The role involves monitoring customer journeys, APIs, and incident management, requiring 5+ years of experience... 

    ClarityPay Program Services, LLC

    New York, NY
    13 hours ago
  •  ...Komodor, a remote-first company, is seeking a Solutions Engineer to connect customer business initiatives to the Komodor platform, understanding developers, DevOps and Incident response teams working with Kubernetes. You will identify customer pain points and communicate... 
    Remote work

    Komodor

    New York, NY
    13 hours ago
  •  ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1... 

    Karsun Solutions

    New York, NY
    13 hours ago
  • $115k - $125k

     ...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help... 
    Work experience placement
    Work at office
    Work from home
    Monday to Friday
    Flexible hours
    Shift work
    Weekend work
    Afternoon shift
    Early shift

    Pico

    New York, NY
    4 days ago
  •  ...Participate in an oncall rotation. Work with teams across the company to ensure we achieve the right balance of developer velocity, reliability and performance, and cost efficiency. What You'll Bring ~5+ years of experience ~ Experience with containerization and... 

    Clay Labs

    New York, NY
    4 days ago
  •  ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client... 

    Karsun Solutions

    New York, NY
    13 hours ago
  •  ...A dynamic fintech company in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability of its fintech products. The role focuses on end-to-end monitoring of customer journeys, incident management, and collaboration with various teams. Candidates... 

    ClarityPay

    New York, NY
    13 hours ago
  •  ...researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At...  ...new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform and customer... 

    Modal Labs

    New York, NY
    13 hours ago
  • $225k - $325k

     ...What you'll do day-to-day Ensure the scalability, reliability, and observability of our systems to maintain and improve the firm's core infrastructure environment. Lead a range of engineering projects, from developing proprietary platforms for configuration... 
    Hourly pay

    D. E. Shaw & Co.

    New York, NY
    2 days ago
  • $100k - $250k

     ...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial...  ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics... 
    Local area

    Kalshi Inc

    New York, NY
    4 days ago
  •  ...No one coasts. If you're driven by impact, pace, and raising the bar. This is the place. The role As a Senior Site Reliability Engineer you'll join the founding SRE team at our new NYC engineering hub, sitting within Foundations. You'll own critical services... 
    Work at office

    Legora

    New York, NY
    4 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern... 
    Local area

    2T Consulting

    New York, NY
    1 day ago
  • $115k - $160k

     ...with Barclays to connect them with exceptional professionals for this role. Embark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear –to redefine the future of banking and help craft innovative... 
    Hourly pay
    Work at office

    Barclays

    New York, NY
    2 days ago
  • $160k - $180k

     ...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience... 

    Socure Inc

    New York, NY
    13 hours ago
  •  ...Participate in an oncall rotation. Work with teams across the company to ensure we achieve the right balance of developer velocity, reliability and performance, and cost efficiency. What You'll Bring ~5+ years of experience ~ Experience with containerization... 

    CLAY

    New York, NY
    3 days ago
  •  ...Federal Reserve Bank of New York is seeking an experienced Cloud AWS Support Reliability Engineer (SRE) to build and maintain scalable AWS infrastructure and CI/CD pipelines. The role emphasizes observability, security, and resilience across enterprise cloud platforms... 

    Federal Reserve Bank (NY)

    New York, NY
    1 day ago
  • $93k - $160k

     ...Palantir Technologies is seeking a Site Reliability Operations Analyst in New York, NY. In this role, you will streamline workflows and reduce friction in deployments. Your responsibilities include supporting deployments, removing roadblocks, and managing multiple challenges... 

    Dormont Manufacturing Company

    New York, NY
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!