Site Reliability Engineer (SRE)
Longfinch Technologies
Overview
We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.
The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.
Key Responsibilities
- Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
- Optimize, and support highly available VDI environments on Hyper-V.
- Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
- Disaster recovery, backup, patch management, and business continuity strategies.
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
- Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
- Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
- Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
- Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
- Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
- Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
- Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
- Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
- Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.
Experience & Qualifications
- 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
- Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
- Proven experience implementing automation to reduce operational overhead and improve service reliability.
- Experience supporting enterprise private cloud and VDI environments.
- Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
- Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
- Experience in Banking or Financial Services environments is advantageous.
Preferred Skills
- Windows Server 2016/2019/2022 administration.
- Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
- Exposure to hybrid cloud and private cloud platforms.
- Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
- Experience supporting enterprise VDI environments.
- Understanding of ITIL Incident, Problem, Change, and Release Management.
- Experience working in regulated industries such as Banking or Financial Services.
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Suggested
- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...SuggestedLocal area
- ...SRE/DevOps EngineerWe are looking for a highly technical, hands-on engineer to join a critical engineering organization supporting enterprise infrastructure and platform... ...software delivery through DevOps and Site Reliability Engineering best practices.Support Infrastructure...Suggested
- ...PVH (Tommy Hilfiger/Calvin Klein) seeks a Senior Software Engineer to own the reliability and performance of our Kubernetes-based data platform across multi-region deployments. You will design scalable infrastructure, optimize deployment pipelines, and strengthen security...Suggested
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively... ...using methodologies & tools such as XP, Lean, DevSecOps, SRE, ADO, GitHub, SonarQube, MLflow, and agentic AI frameworks (...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week$140k - $170k
...experience) Eligibility: U.S. citizenship required (customer badging requirement); no security clearance required The Role As a Site Reliability Engineer on Blitzy's Public Sector team, you will be the backbone of our platform's reliability and operational excellence for a...Remote work- ...meaningful products that make a real impact on children's education and literacy. About the Role We're looking for a Senior Site Reliability Engineer to drive the stability, observability, and reliability of Epic's platform as we grow. You are an experienced engineer who...Remote work
$153k - $204k
...key part in ensuring the availability, reliability, and scalability of one of the industry’... ...operational requirements and customer SLAs.Engineer for resiliency, implementing best... ...of experience in production engineering, SRE, or large-scale infrastructure/platform...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based...Full timeTemporary workWork experience placementRemote work$227k - $303k
...developer-facing capabilities that enable every engineer at CoreWeave to build and ship software... ..., heterogeneous workloads, and the reliability demands of an infrastructure platform... ....Partner with infrastructure, security, SRE, and product engineering teams to ensure...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$153k - $204k
...Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at .What You'll Do:The Systems Engineering team owns the host software stack that turns a freshly provisioned bare-metal machine into a healthy Kubernetes worker — the OS...Permanent employmentFull timeTemporary workCasual workLive inWork at officeFlexible hours$165k - $242k
...and scale CoreWeave's Platform Security engineering function, owning how security is designed... ...Infrastructure, Platform Engineering, SRE, and other security teams to ensure platform... ...for extreme performance, scale, and reliability, supporting frontier AI development across...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$207k - $275k
...CRWV) in March 2025. Learn more at .What You’ll DoThe Endpoint Engineering team at CoreWeave manages end-user compute for all employees... ...and configuration management to keep the environment secure and reliable.We're expanding that scope with a cloud-based VDI offering (...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours- ...Crown difference. Job Description Crown seeks Software Engineers to support development of next generation global positioning... ...systems and inertial navigation systems. Work may be performed on-site (preferred) in Pine Brook, New Jersey, or at Crown's...Full timeFor contractors
$165k - $242k
...Learn more at .What You’ll Do:We are seeking a Senior Platform Engineer to join our Kubernetes Infrastructure team. This role involves... ...architecture and deployment.About the role:Champion reliability initiatives for Kubernetes application deployments: Advocate for...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$153k - $204k
...in March 2025. Learn more at .What You'll Do:The Cloud Platform API team builds the API platform and the tooling that every other engineering team at CoreWeave uses to ship and consume APIs. This role sits with the group that owns CoreWeave's Go framework for building...Permanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours$182k - $242k
...Do:The Box Office Platform team sits within CoreWeave’s Fleet Engineering Organization, which is responsible for the automated provisioning... ...node repairs, and vendor integrations into a cohesive, highly reliable fleet management engine.About the role:As a Senior Software...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$139k - $242k
...lightweight virtualization, GPU infrastructure, and Linux systems engineering. We partner closely with security, platform, and GPU... ...knowledge.Skilled at diagnosing and resolving complex performance, reliability, or isolation issues across containers, VMs, and...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$109k - $204k
...team owns the platform, services, and systems that enable every engineer at CoreWeave to build and ship software faster, more safely,... ...for artifact publishing, retrieval, and version management that reliably serve engineering teams at scale.Identify inefficiencies and reliability...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$165k - $242k
...CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at .About the RoleWe're looking for a software engineer to join our Source Control and Governance team within our Developer Experience group. In this role, you'll design and build the...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...in March 2025. Learn more at .What You'll Do:The Platform & Infrastructure Engineering team sits at the core of our Data Infrastructure organization, responsible for the availability, reliability, scalability, and security of the company's data platform. We build and...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$165k - $242k
...team owns the platform, services, and systems that enable every engineer at CoreWeave to build and ship software faster, more safely,... ...The work you do will directly shape developer velocity, system reliability, and the agentic developer experience across a rapidly scaling...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$165k - $242k
...will work on and extend our custom job scheduler, improving reliability, observability, and execution guarantees for distributed workloads... .... You will collaborate closely with Product and various Engineering teams to design systems that are reliable, scalable, and maintainable...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$139k - $204k
...traded company (Nasdaq: CRWV) in March 2025. Learn more at .What You’ll Do:CoreWeave is seeking a highly skilled and motivated Systems Engineer to join our People Systems team. In this role, you will be a key contributor to the design, implementation, maintenance, and...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$165k - $242k
...CoreWeave is seeking a highly skilled and motivated Systems Kernel Engineer to join our HAVOCK Team, reporting into the Manager of Systems... ...fixes and features that improves the performance and reliability of our stack.This position is ideal for someone who thrives in...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$153k - $204k
...Learn more at .What You’ll DoThe AI Dev Platform Tooling team builds and operates the SaaS platform that lets our engineers ship software quickly, reliably, and safely. We own the deployment automation, infrastructure provisioning and monitoring, and release systems — built...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at .What You'll Do:The Systems Engineering team owns the Linux kernel and host software stack underneath one of the largest GPU fleets in the world. When something breaks...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$109k - $160k
...the Kubernetes platforms running CoreWeave’s GPU workloads are reliable, fault-tolerant, and easy to operate, providing deep insights and... ...for users and internal teams alike.About the role:As an Engineer on the Kubernetes Core Interfaces Team, you’ll design and implement...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...company (Nasdaq: CRWV) in March 2025. Learn more at .What You'll Do:CoreWeave is seeking a highly skilled and motivated Senior Systems Engineer, Legal Systems to build and scale our contract lifecycle management (CLM) and legal technology ecosystem, with a primary focus on...Permanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours$250k - $275k
...is seeking a Vice President of Software Engineering to lead the design, development, and modernization... ...for CI/CD, code quality, observability, reliability, and operational excellence.Ensure... ....Strong understanding of DevOps, SRE, and modern observability practices.Experience...Full timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- on-site clinical research associate (traveling/remote) Caldwell, NJ
- site reliability engineer
- lead site reliability engineer
- junior site reliability engineer
- site reliability engineering manager
- site reliability engineer remote
- site reliability engineer sre
- site paramedic
- wordpress website builder
- website content developer


