Site Reliability Engineer (SRE)
Longfinch Technologies
Overview
We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.
The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.
Key Responsibilities
- Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
- Optimize, and support highly available VDI environments on Hyper-V.
- Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
- Disaster recovery, backup, patch management, and business continuity strategies.
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
- Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
- Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
- Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
- Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
- Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
- Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
- Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
- Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
- Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.
Experience & Qualifications
- 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
- Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
- Proven experience implementing automation to reduce operational overhead and improve service reliability.
- Experience supporting enterprise private cloud and VDI environments.
- Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
- Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
- Experience in Banking or Financial Services environments is advantageous.
Preferred Skills
- Windows Server 2016/2019/2022 administration.
- Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
- Exposure to hybrid cloud and private cloud platforms.
- Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
- Experience supporting enterprise VDI environments.
- Understanding of ITIL Incident, Problem, Change, and Release Management.
- Experience working in regulated industries such as Banking or Financial Services.
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Suggested
- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...SuggestedLocal area
- ...First Due seeks a Director of Platform Engineering to lead DevOps, SRE, and DBRE, turning fragmented practices into a centralized, disciplined... ...hands-on depth and executive presence to represent platform reliability to the ELT. Reporting to SVP of Engineering, this leader...Suggested
$83.54k - $137.24k
...progress and enhancing lives by providing reliable, high-speed connectivity solutions that... ...! Job Summary The Role DNS Engineer – SRE is a high-impact role responsible for the... ...views infrastructure through the lens of Site Reliability Engineering (SRE)...SuggestedLocal areaRemote work- LiveRamp, a data collaboration platform leader, is seeking a Senior Site Reliability Engineer in San Francisco with 5+ years of SRE/DevOps experience. The role focuses on deployments, 24/7 support across regions, and establishing SRE best practices. Strong skills in Terraform...Suggested
$253k - $336k
...TEAM: CorpTech Platform is the internal engineering force multiplier behind Anduril's... ...products. ABOUT THE JOB: The Director of Site Reliability Engineering owns the reliability system... ...manufacturing operations. This role leads the SRE organization and partners with software...Full timeWork experience placement$174k - $252k
Senior Software Engineer, Site Reliability Engineering corporate_fare Google place Seattle, WA, USA ; Kirkland, WA, USA Mid Experience driving progress... ...Engineering. About the job Site Reliability Engineering (SRE) is what you get when you treat operations as if it’s a...Temporary work- Google is seeking a Senior Software Engineer in Site Reliability Engineering to strengthen the reliability of Google's public services from the Seattle/Kirkland area. You will design, build, and operate scalable systems with emphasis on availability and performance. The...
- ...help healthcare organizations streamline operations through reliable, customizable technology that improves efficiency and... ...the Medgen environment . We are looking for a hands-on Site Reliability Engineer / Infrastructure Engineer to own the reliability, security...Full timeWork at officeTrial period
- ...help healthcare organizations streamline operations through reliable, customizable technology that improves efficiency and... ...the Labgen environment . We are looking for a hands-on Site Reliability Engineer / Infrastructure Engineer to own the reliability, security...Full timeContract workWork at officeTrial period
- Google is seeking a Software Engineering Manager for Site Reliability Engineering in Durham, NC. You will lead a team responsible for uptime, availability, and performance of key services, and drive automation to prevent problems. The role requires deep experience in programming...
$88k - $168k
Do you bring deep full-stack software engineering experience and the judgment to identify, design and deliver the right solutions? Are you... ...the globe. Your Mission As a Principal Software Engineer in our SRE & Architecture team within Tech, reporting to the Head of...Permanent employmentContract workTemporary workWork at officeLocal areaWork from homeWorldwideWeekend work$132.23k - $176.31k
...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...Full timeTemporary workRemote work- Infosys Limited is seeking a Technology Consultant 1 in the US to support, design, and document software solutions. You will work across SDLC lifecycle, develop code, and help troubleshoot production issues while collaborating with Dev teams. This role emphasizes agile ...
- ...Mount Thor, Inc. is seeking an on-site engineer to design and improve the data center systems behind our compute fleet. You will own the technical direction for data center architecture and operations, define standards for power, cooling, cabling, and network integration...
- 3Core Systems, Inc. is seeking a Senior Splunk/Data Management Engineer to design, implement, and optimize Splunk infrastructure to support scalable, robust data solutions. The role emphasizes SRE practices, incident response, and cross-functional leadership across multiple...
- Ford is seeking a Director of Cloud SRE to lead a global team focused on federating core SRE principles across hybrid environments... ...strategy, architecture, and roadmaps for internal observability and reliability tooling, spanning telemetry standards, pipeline integration,...
- ...We are seeking a Secure Enterprise Browser DevOps Engineer with strong production support and reliability engineering experience. The role focuses on maintaining... ...for enterprise environments. Collaborate with SRE, security, engineering, and incident response teams...Permanent employment
- ...best work together.What's the role?We are looking for a Software Engineer to help improve Etsy's development workflows and environments.... ...toolingWe work closely with other infrastructure, product, and SRE teams to ensure consistency between development and productionWe...Full timeWork at officeLocal areaRemote workVisa sponsorship
- MillTech is seeking an Application Support Engineer in New York City. This in-office role requires 3 days per week on-site to support platform reliability and scale. You will own incident response, automate workflows, and collaborate with engineering and business teams...Work at office3 days per week
- Senior Software Systems Engineer (Storage) - remote in the US Important: if an employer asks... ...-directed in diagnosing performance and reliability issues end to end, set operational... ...qualifications: 7+ years of experience in SRE or hardware/storage infrastructure operations...Remote jobLocal area
- ...looking for entry-level software programmers, Java full stack developers, Python/Java developers, data analysts/data scientists, data engineers, machine learning engineers for full time positions with clients. Who should apply? Recent computer science/engineering/...Full time
- ...Senior Software Engineer – Frontend We are seeking a Frontend Engineer to design and scale AI-powered applications that automate complex professional workflows. You will work closely with the leadership team and domain experts to deploy systems used by specialized...H1bRelocationVisa sponsorship
- ...issues. Collaborate with architects, business stakeholders, and engineering teams to deliver high-quality solutions. Ensure applications meet performance, scalability, security, and reliability requirements. Contribute to technical strategy, system design, and...
$15 per hour
Senior Software Engineer, Wikidata Platform Summary The Wikimedia Foundation is seeking... ...grade services while ensuring performance, reliability, and maintainability. Working closely... ...sources across Wikimedia Collaborate with SRE, data engineers, and product teams to...Permanent employmentFor contractorsCurrently hiringLocal areaRemote work- ...and test solutions for healthcare data exchanges, focusing on EDI (x12 HIPAA) and related formats such as HL7/FHIR. Collaboration with SRE and cross‑functional teams is essential. The role requires 5+ years of software development experience, backend work with Java or...
- Senior Software Engineer, Platform FullTime Professional Round Rock, TX, US 16 days ago Requisition... ...architecture and data model into reliable production services spanning object... ...quality assurance, and production support (SRE-lite) Implement platform observability,...Full timeContract work
- ...enterprise secrets management solutions.(Preferred based on role requirements; not explicitly present in the resume.) Site Reliability Engineering (SRE) concepts including observability, reliability, and operational excellence. Financial Services, Regulatory Reporting,...Full timeTemporary workRelocation
- ...fundamentals and problem-solving skills (such as data structures, computational algorithms, and operating systems)0-1+ years of data engineering experienceBenefits:On job technical supportE-verifiedFull time positionCandidates who are missing the required skills might be...Immediate start
- ...Automation Engineer — AI Agents & IntegrationsBold Business is a U.S.-based, AI-first company building automation systems, AI agents and... ...understand their workflows, confirm requirements and deliver reliable solutions. We want someone who takes ownership while communicating...Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- on-site clinical research associate (traveling/remote) Jamaica, NY
- construction site safety Jamaica, NY
- site reliability engineer
- site reliability engineering manager
- junior site reliability engineer
- site reliability engineer sre
- site reliability engineer remote
- lead site reliability engineer
- savannah river site
- site inspector



