Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Kontakt.io

Lead Site Reliability Engineer

We're changing the way hospitals operate by combining proprietary hardware, AI-powered intelligence, and deep integrations with the systems hospitals already rely on. We're creating a real-time understanding of hospital operations that software alone can't deliver. This intelligence powers the execution layer hospitals have been missing – helping care teams make smarter decisions and deliver better patient care.

If you're excited to solve hard problems, work with a team of builders, and help hospitals deliver better care every day, we'd love to meet you!

What You'll Do
  • Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.

  • Design and implement self-healing, fault-tolerant systems to prevent failures before they happen.

  • Define SLIs, SLOs, and SLAs, ensuring proactive performance monitoring and incident resolution.

  • Architect and manage scalable cloud infrastructure (AWS) for massive real-time data processing.

  • Optimize containerized environments (Kubernetes, Docker) to support multi-region deployments.

  • Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.

  • Build and refine a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.

  • Lead incident response and on-call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR).

  • Conduct blameless postmortems and continuously improve system resilience.

  • Reduce manual intervention through automated deployment, scaling, and failover mechanisms.

  • Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.

  • Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.

  • Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.

  • Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.

What You Have
  • 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.

  • Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.

  • Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.

  • Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.

  • Hands-on experience with incident management, postmortems, and building resilient systems.

  • Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).

  • A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team.

  • Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2).

Bonus Points If You Have:
  • Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.

  • Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines.

  • Prior experience leading on-call rotations and major incident management processes.

Logistics, Perks & Benefits
  • Built for collaboration - our team a hybrid schedule of 3 days/week minimum from our New York City office

  • Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors

  • Full health, dental, and vision coverage, a 401k, paid time off, paid parental leave and all the tools you need to do your best work

  • Autonomy to solve meaningful problems with work that ships quickly and makes a difference

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in New York, NY vacancy
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity... 
    Suggested
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    New York, NY
    4 days ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank...  ...tooling to reduce manual steps and speed repeatable triage.Lead L1/L2 production support using SRE practices: quickly... 
    Suggested
    Shift work

    JP Morgan Chase

    New York, NY
    12 hours ago
  • $120k - $200k

     ...DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal...  ...cause analysis, and postmortem processes. Lead cross-team coordination during major incidents...  ...testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Suggested
    Overseas

    Comrise

    New York, NY
    1 day ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing...  ...within an IT organization is a plusPrior experience leading a technical team preferredThe estimated base salary range... 
    Suggested
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    5 days ago
  •  ...developers on how to make things better. We collectively strive to build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS & Karpenter Management: Manage, upgrade, and autoscale... 
    Suggested
    For contractors

    Varo Money

    New York, NY
    1 day ago
  • $167.7k - $245.2k

     ...assurance insights within Cisco’s leading Networking, Security,...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    3 days ago
  • $158.5k - $172k

     ...velocity energy of a powerhouse startup.As a leading U.S. ordering and delivery marketplace,...  ....About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will...  ...high-impact position driving continuous reliability, deep system optimization, and automation... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    4 days ago
  • $141k - $216.6k

     ...building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational...  ...of ambiguous technical problems with minimal direction.Leading operational improvements rather than simply executing assigned... 
    Work experience placement
    Work at office

    Axon

    New York, NY
    2 days ago
  • $110k - $120k

    As a leading financial services and healthcare technology company based on revenue, SS&C is headquartered in Windsor, Connecticut...  ...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    1 day ago
  • $139k - $257.55k

     ...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,...  ...productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express,... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    1 day ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    1 day ago
  • $150k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible...  ...Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the...  ...into durable engineering improvements.Lead incident response for latency-sensitive... 
    Full time
    Temporary work
    Part time

    Goldman Sachs

    New York, NY
    3 days ago
  •  ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership... 

    JP Morgan Chase

    New York, NY
    2 days ago
  • Who are we?Cohere is the leading security-first enterprise AI company...  ...is a team of researchers, engineers, designers, and more, who are...  ...high-performance, scalable and reliable machine learning systems? Do...  ...? We are looking for a Site Reliability Engineer to join... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    4 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    3 days ago
  • $131k - $164k

    Position OverviewWe are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation...  ...they need to drive greater impact and accountability - to lead with purpose. Our employees are passionate, smart, and... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Diligent

    New York, NY
    4 days ago
  • $150k - $190k

    Senior Site Reliability Engineer, VPAt Morgan Stanley, we advise, originate, trade, manage and distribute capital for governments, institutions...  ..., and always do so with a standard of excellence. We are a leading global financial services firm that conducts its business through... 
    Temporary work
    Worldwide
    Flexible hours
    Weekend work

    Morgan Stanley

    New York, NY
    4 days ago
  • $130k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible...  ...your journey here.Securities Frontline Site Reliability Engineers (SREs) play a critical role...  ...serve to grow. Founded in 1869, we are a leading global investment banking, securities... 
    Full time
    Temporary work
    Part time
    Immediate start

    Goldman Sachs

    New York, NY
    3 days ago
  • $195k - $275k

    Morgan Stanley is a leading global financial services firm providing a wide range of investment banking, securities, investment management...  ...& Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift,... 
    Temporary work
    Work at office
    Worldwide
    Night shift

    Morgan Stanley

    New York, NY
    2 days ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency... 
    Local area
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    3 days ago
  •  ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your... 
    Full time
    Work visa

    Syntricate Technologies

    New York, NY
    3 days ago
  •  ...Site Reliability Engineer TXSE is building the next-generation exchange infrastructure to support transparent, efficient, and resilient capital markets. With SEC approval and $275MM in funding, we are currently hiring a Site Reliability Engineer to help with a greenfield... 
    Currently hiring

    TXSE

    New York, NY
    4 days ago
  •  ...and Antler, we empower CISOs to proactively manage human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our... 
    Full time
    Work at office

    Dune Security

    New York, NY
    2 days ago
  •  ...Triomics Backend Engineer Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained staff manually reading unstructured patient records - pathology reports, clinical notes, genomic panels - to power workflows like... 
    Day shift

    Triomics

    New York, NY
    1 day ago
  • $86k - $105k

     ...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best...  ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    New York, NY
    5 days ago
  •  ...shifting towards Linux – (70% Windows, 30% Linux). Remote access technology protocols are a plus. Job Description Site Reliability Engineer Periodic updates and maintenance of Windows-based golden image for ESX & AWS. Patching of software, systems,... 
    Remote work
    Shift work

    Omni Inclusive

    New York, NY
    3 days ago
  • $150k - $160k

    Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join...  ...fast, resilient, and optimized at the edge.Responsibilities:Lead development team initiatives to:Architect and maintain Cloudflare... 
    Work at office
    Local area

    Haymarket Media Group

    New York, NY
    4 days ago
  •  ...Sr. Site Reliability Engineer (SRE) New York City, NY - LOCALS ONLY Hybrid, 3 days 6-Month Contract 10-15 years Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production... 
    Contract work
    Local area

    RIT Solutions

    New York, NY
    2 days ago
  •  ...SRE Engineer Location: New York, NY, USA Experience: 8-12 Years Client: Amex Job Description: This is an SRE role supporting the B2B and Core Services. This is not a DevOps role, strictly need an SRE Engineer, who has great analytical skills and is a good... 

    IVID TEK INC

    New York, NY
    22 hours ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Local area

    E-Solutions

    New York, NY
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!