Lead Site Reliability Engineer
Kontakt.io
Lead Site Reliability Engineer
We're changing the way hospitals operate by combining proprietary hardware, AI-powered intelligence, and deep integrations with the systems hospitals already rely on. We're creating a real-time understanding of hospital operations that software alone can't deliver. This intelligence powers the execution layer hospitals have been missing – helping care teams make smarter decisions and deliver better patient care.
If you're excited to solve hard problems, work with a team of builders, and help hospitals deliver better care every day, we'd love to meet you!
What You'll Do
Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.
Design and implement self-healing, fault-tolerant systems to prevent failures before they happen.
Define SLIs, SLOs, and SLAs, ensuring proactive performance monitoring and incident resolution.
Architect and manage scalable cloud infrastructure (AWS) for massive real-time data processing.
Optimize containerized environments (Kubernetes, Docker) to support multi-region deployments.
Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.
Build and refine a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.
Lead incident response and on-call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR).
Conduct blameless postmortems and continuously improve system resilience.
Reduce manual intervention through automated deployment, scaling, and failover mechanisms.
Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.
Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.
Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.
What You Have
10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.
Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.
Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.
Hands-on experience with incident management, postmortems, and building resilient systems.
Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).
A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team.
Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2).
Bonus Points If You Have:
Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.
Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines.
Prior experience leading on-call rotations and major incident management processes.
Logistics, Perks & Benefits
Built for collaboration - our team a hybrid schedule of 3 days/week minimum from our New York City office
Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors
Full health, dental, and vision coverage, a 401k, paid time off, paid parental leave and all the tools you need to do your best work
Autonomy to solve meaningful problems with work that ships quickly and makes a difference
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank... ...tooling to reduce manual steps and speed repeatable triage.Lead L1/L2 production support using SRE practices: quickly...SuggestedShift work
$120k - $200k
...DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal... ...cause analysis, and postmortem processes. Lead cross-team coordination during major incidents... ...testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)SuggestedOverseas$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing... ...within an IT organization is a plusPrior experience leading a technical team preferredThe estimated base salary range...SuggestedWork at officeLocal areaImmediate start- ...developers on how to make things better. We collectively strive to build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.ResponsibilitiesEKS & Karpenter Management: Manage, upgrade, and autoscale...SuggestedFor contractors
$167.7k - $245.2k
...assurance insights within Cisco’s leading Networking, Security,... ...effective.We’re looking for talented engineers with a software or operations... ...teams to ensure the reliability, performance and security of... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$158.5k - $172k
...velocity energy of a powerhouse startup.As a leading U.S. ordering and delivery marketplace,... ....About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and automation...Full timeTemporary workWork at officeFlexible hours3 days per week$141k - $216.6k
...building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational... ...of ambiguous technical problems with minimal direction.Leading operational improvements rather than simply executing assigned...Work experience placementWork at office$110k - $120k
As a leading financial services and healthcare technology company based on revenue, SS&C is headquartered in Windsor, Connecticut... ...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading...Ongoing contractFull timeCasual workRemote workFlexible hours$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...productivity and personalized customer experiences. Adobe’s industry-leading offerings including Adobe Acrobat Studio, Adobe Express,...Full timeTemporary workLocal areaRemote workWorldwide$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$150k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible... ...Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the... ...into durable engineering improvements.Lead incident response for latency-sensitive...Full timeTemporary workPart time- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership...
- Who are we?Cohere is the leading security-first enterprise AI company... ...is a team of researchers, engineers, designers, and more, who are... ...high-performance, scalable and reliable machine learning systems? Do... ...? We are looking for a Site Reliability Engineer to join...Full timeWork experience placementWork at officeLocal areaRemote workHome office
$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$131k - $164k
Position OverviewWe are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation... ...they need to drive greater impact and accountability - to lead with purpose. Our employees are passionate, smart, and...Work at officeLocal areaVisa sponsorshipFlexible hours$150k - $190k
Senior Site Reliability Engineer, VPAt Morgan Stanley, we advise, originate, trade, manage and distribute capital for governments, institutions... ..., and always do so with a standard of excellence. We are a leading global financial services firm that conducts its business through...Temporary workWorldwideFlexible hoursWeekend work$130k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible... ...your journey here.Securities Frontline Site Reliability Engineers (SREs) play a critical role... ...serve to grow. Founded in 1869, we are a leading global investment banking, securities...Full timeTemporary workPart timeImmediate start$195k - $275k
Morgan Stanley is a leading global financial services firm providing a wide range of investment banking, securities, investment management... ...& Release Management, and the Chief Operating Office.The Reliability Operations (RO) within WMT is responsible for providing swift,...Temporary workWork at officeWorldwideNight shift$111k - $218k
...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency...Local areaWorldwideFlexible hours- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
- ...Site Reliability Engineer TXSE is building the next-generation exchange infrastructure to support transparent, efficient, and resilient capital markets. With SEC approval and $275MM in funding, we are currently hiring a Site Reliability Engineer to help with a greenfield...Currently hiring
- ...and Antler, we empower CISOs to proactively manage human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our...Full timeWork at office
- ...Triomics Backend Engineer Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained staff manually reading unstructured patient records - pathology reports, clinical notes, genomic panels - to power workflows like...Day shift
$86k - $105k
...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best... ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...shifting towards Linux – (70% Windows, 30% Linux). Remote access technology protocols are a plus. Job Description Site Reliability Engineer Periodic updates and maintenance of Windows-based golden image for ESX & AWS. Patching of software, systems,...Remote workShift work
$150k - $160k
Front-End & AdTech Site Reliability Engineer (SRE)Haymarket Media, Inc. is seeking a Front-End & AdTech Site Reliability Engineer (SRE) to join... ...fast, resilient, and optimized at the edge.Responsibilities:Lead development team initiatives to:Architect and maintain Cloudflare...Work at officeLocal area- ...Sr. Site Reliability Engineer (SRE) New York City, NY - LOCALS ONLY Hybrid, 3 days 6-Month Contract 10-15 years Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production...Contract workLocal area
- ...SRE Engineer Location: New York, NY, USA Experience: 8-12 Years Client: Amex Job Description: This is an SRE role supporting the B2B and Core Services. This is not a DevOps role, strictly need an SRE Engineer, who has great analytical skills and is a good...
- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead backend developer New York, NY
- lead app. developer New York, NY
- lead web developer New York, NY
- lead algorithm engineer New York, NY
- lead infrastructure engineer New York, NY
- lead network engineer New York, NY
- lead engineer New York, NY
- lead operating engineer New York, NY
- lead system engineer New York, NY
- site reliability engineer remote New York, NY

