Site Reliability Engineer (SRE)
Longfinch Technologies
Overview
We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.
The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.
Key Responsibilities
- Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
- Optimize, and support highly available VDI environments on Hyper-V.
- Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
- Disaster recovery, backup, patch management, and business continuity strategies.
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
- Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
- Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
- Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
- Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
- Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
- Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
- Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
- Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
- Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.
Experience & Qualifications
- 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
- Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
- Proven experience implementing automation to reduce operational overhead and improve service reliability.
- Experience supporting enterprise private cloud and VDI environments.
- Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
- Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
- Experience in Banking or Financial Services environments is advantageous.
Preferred Skills
- Windows Server 2016/2019/2022 administration.
- Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
- Exposure to hybrid cloud and private cloud platforms.
- Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
- Experience supporting enterprise VDI environments.
- Understanding of ITIL Incident, Problem, Change, and Release Management.
- Experience working in regulated industries such as Banking or Financial Services.
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Suggested
- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...SuggestedLocal area
- ...SRE/DevOps EngineerWe are looking for a highly technical, hands-on engineer to join a critical engineering organization supporting enterprise infrastructure and platform... ...software delivery through DevOps and Site Reliability Engineering best practices.Support Infrastructure...Suggested
- ...PVH (Tommy Hilfiger/Calvin Klein) seeks a Senior Software Engineer to own the reliability and performance of our Kubernetes-based data platform across multi-region deployments. You will design scalable infrastructure, optimize deployment pipelines, and strengthen security...Suggested
- ...SRE Senior EngineerLocation: Union, NJRate: DOE $/hr. on w2 onlyPosition Type: contractResponsibilities:Engage at all levels across the organizationDesign, build, test, and size systemsAutomating monitoring, alerting, and escalationTroubleshoot issues across the entire...Suggested
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively... ...using methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks (...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...meaningful products that make a real impact on children's education and literacy. About the Role We're looking for a Senior Site Reliability Engineer to drive the stability, observability, and reliability of Epic's platform as we grow. You are an experienced engineer who...Remote work
$140k - $170k
...experience) Eligibility: U.S. citizenship required (customer badging requirement); no security clearance required The Role As a Site Reliability Engineer on Blitzy's Public Sector team, you will be the backbone of our platform's reliability and operational excellence for a...Remote work$153k - $204k
...key part in ensuring the availability, reliability, and scalability of one of the industry’... ...operational requirements and customer SLAs.Engineer for resiliency, implementing best... ...of experience in production engineering, SRE, or large-scale infrastructure/platform...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based...Full timeTemporary workWork experience placementRemote work$227k - $303k
...developer-facing capabilities that enable every engineer at CoreWeave to build and ship software... ..., heterogeneous workloads, and the reliability demands of an infrastructure platform... ....Partner with infrastructure, security, SRE, and product engineering teams to ensure...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$207k - $275k
...CRWV) in March 2025. Learn more at .What You’ll DoThe Endpoint Engineering team at CoreWeave manages end-user compute for all employees... ...and configuration management to keep the environment secure and reliable.We're expanding that scope with a cloud-based VDI offering (...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$165k - $242k
...and scale CoreWeave's Platform Security engineering function, owning how security is designed... ...Infrastructure, Platform Engineering, SRE, and other security teams to ensure platform... ...for extreme performance, scale, and reliability, supporting frontier AI development across...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$153k - $204k
...Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at .What You'll Do:The Systems Engineering team owns the host software stack that turns a freshly provisioned bare-metal machine into a healthy Kubernetes worker — the OS...Permanent employmentFull timeTemporary workCasual workLive inWork at officeFlexible hours$153k - $204k
...in March 2025. Learn more at .What You'll Do:The Cloud Platform API team builds the API platform and the tooling that every other engineering team at CoreWeave uses to ship and consume APIs. This role sits with the group that owns CoreWeave's Go framework for building...Permanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours$165k - $242k
...Learn more at .What You’ll Do:We are seeking a Senior Platform Engineer to join our Kubernetes Infrastructure team. This role involves... ...architecture and deployment.About the role:Champion reliability initiatives for Kubernetes application deployments: Advocate for...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...Do:The Box Office Platform team sits within CoreWeave’s Fleet Engineering Organization, which is responsible for the automated provisioning... ...node repairs, and vendor integrations into a cohesive, highly reliable fleet management engine.About the role:As a Senior Software...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...We're building the next generation of ADP technologies. As a Full Stack Application Developer, you'll work in a team of software engineers within the Insurance Services for Small Business Solutions Technology team.Like what you see?Apply now!Learn more about ADP at tech...
$139k - $242k
...lightweight virtualization, GPU infrastructure, and Linux systems engineering. We partner closely with security, platform, and GPU... ...knowledge.Skilled at diagnosing and resolving complex performance, reliability, or isolation issues across containers, VMs, and...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$165k - $242k
...team owns the platform, services, and systems that enable every engineer at CoreWeave to build and ship software faster, more safely,... ...The work you do will directly shape developer velocity, system reliability, and the agentic developer experience across a rapidly scaling...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...in March 2025. Learn more at .What You'll Do:The Platform & Infrastructure Engineering team sits at the core of our Data Infrastructure organization, responsible for the availability, reliability, scalability, and security of the company's data platform. We build and...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$109k - $204k
...team owns the platform, services, and systems that enable every engineer at CoreWeave to build and ship software faster, more safely,... ...for artifact publishing, retrieval, and version management that reliably serve engineering teams at scale.Identify inefficiencies and reliability...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$94.4k - $198.2k
...Type: RegularPercentage of Travel Required: Up to 10%Type of Travel: Continental US* * *The Opportunity:We are seeking a Software Engineer to design and develop advanced cyber capabilities that strengthen national security. In this role, you'll work across real‑time embedded...Contract workWork experience placementFlexible hours$165k - $242k
...CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at .About the RoleWe're looking for a software engineer to join our Source Control and Governance team within our Developer Experience group. In this role, you'll design and build the...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$124k - $280k
...ApplicableSpecialismData, Analytics & AIManagement LevelSenior ManagerJob Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design and develop robust data solutions for clients. They play a...Full timeH1b$99k - $232k
...ApplicableSpecialismData, Analytics & AIManagement LevelManagerJob Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design and develop robust data solutions for clients. They play a crucial...Full timeH1b- ...Job Description Job Description Hiring: Release Train Engineer (RTE) – SAFe Agile Location: Bedminster, NJ (Hybrid – 3 Days/Week Onsite) Experience: 11+ Years No CPT/OPT We are looking for an experienced Release Train Engineer (RTE) to lead Agile Release...Work from homeFlexible hours3 days per week
- ...future as we are, join our team.KPMG is currently seeking a Lead Engineer, Network Security Platform to join our Digital Nexus technology... ...benefits can be found towards the bottom of our KPMG US Careers site at Benefits & How We Work.Follow this link to obtain salary ranges...H1bLocal area
- ...find new areas of inspiration and expand your capabilities, then consider a career in Advisory.KPMG is currently seeking a SAP Sales Engineer with Order to Cash expertise to join our Alliances Organization.Responsibilities:Develop and deliver compelling solution...Contract workH1bLocal area
$165k - $242k
...will work on and extend our custom job scheduler, improving reliability, observability, and execution guarantees for distributed workloads... .... You will collaborate closely with Product and various Engineering teams to design systems that are reliable, scalable, and maintainable...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- on-site clinical research associate (traveling/remote) Livingston, NJ
- site reliability engineer
- lead site reliability engineer
- junior site reliability engineer
- site reliability engineering manager
- site reliability engineer remote
- site reliability engineer sre
- site paramedic
- wordpress website builder
- website content developer


