Site Reliability Engineer
2T Consulting
We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.
The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.
Required Technical Skills
- Strong understanding of Site Reliability Engineering principles and operational excellence.
- Experience with infrastructure reliability, service availability, resiliency, and performance optimization.
- Storage Space Direct and failover clustering technical expertise. (Storage Spaces Direct enables you to build highly available, software-defined storage by pooling local disks (SSDs, NVMe drives, and HDDs) across multiple Windows Server nodes in a cluster. Instead of relying on an external SAN, S2D uses the servers' local storage to create a resilient shared storage pool)
- Experience managing production-critical infrastructure environments with high availability requirements.
- Experience with incident management, problem management, RCA, and continuous operational improvement.
- Knowledge of monitoring, observability, alerting, and performance management.
Microsoft Hyper-V (Core Expertise)
- Deep hands-on expertise in Microsoft Hyper-V architecture, deployment, administration, troubleshooting, and optimization.
- Extensive experience in operating enterprise private cloud environments on Hyper-V.
- Strong experience supporting enterprise-scale VDI deployments on Hyper-V.
- Hyper-V Failover Clustering and high-availability architecture.
- Storage integration including SAN, NAS, Storage Spaces Direct (S2D), Cluster Shared Volumes (CSV), and storage optimization.
- Networking within Hyper-V environments including virtual switches, VLANs, NIC Teaming, QoS, and network performance tuning.
- System Center Virtual Machine Manager (SCVMM).
Automation & Platform Engineering
- Strong PowerShell scripting and automation experience.
- Experience automating infrastructure deployment, operational tasks, health checks, and reporting.
- Familiarity with Infrastructure as Code concepts and configuration management.
- Experience developing reusable operational tooling to improve reliability and reduce manual effort.
Preferred Skills
- Windows Server 2016/2019/2022 administration.
- Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
- Exposure to hybrid cloud and private cloud platforms.
- Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
- Experience supporting enterprise VDI environments.
- Understanding of ITIL Incident, Problem, Change, and Release Management.
- Experience working in regulated industries such as Banking or Financial Services.
Experience & Qualifications
- 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
- Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
- Proven experience implementing automation to reduce operational overhead and improve service reliability.
- Experience supporting enterprise private cloud and VDI environments.
- Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
- Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
- Experience in Banking or Financial Services environments is advantageous.
Key Responsibilities
- Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
- Optimize, and support highly available VDI environments on Hyper-V.
- Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
- Disaster recovery, backup, patch management, and business continuity strategies.
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
- Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
- Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
- Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
- Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
- Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
- Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
- Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
- Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
- Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.
- ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern...SuggestedLocal area
- ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise...Suggested
- ...Goldman Sachs is seeking a Vice President in Compliance Engineering SRE for Dallas. The role combines software and systems engineering... ...monitoring, and collaborate with cross-functional teams to deliver reliable, compliant platforms for regulatory risk management. #J-18808-...Suggested
- ...resolution or escalation (ServiceNow/Jira). Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles). Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team. Maintain and update operational runbooks...SuggestedShift workNight shift
$155k - $175k
...Next! Summary We are seeking a highly skilled and experienced Site Reliability Manager to join our team to ensure the reliability,... ...performance of our systems and services. You will lead a team of engineers focusing on three core pillars: Application Reliability, DevSecOps...SuggestedWork experience placementH1bWork at officeLocal area$140k - $150k
WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (...Full timeTemporary workLocal areaRemote workWeekend work- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex...Work at office
- Compliance Engineering, Site Reliability Engineering, Vice President, Dallas location_on Dallas, TX, United States We are Compliance Engineering, a global team of more than 300 engineers and scientists who work on the most complex, mission-critical problems. We build and...Full timeTemporary workWork at office
- Jack Henry & Associates, Inc. is seeking a Senior Site Reliability Engineer to drive modernization across a large-scale hybrid cloud footprint, with emphasis on re-architecting on-prem workloads to Google Cloud Platform. The role involves implementing SRE practices, IaC...
$90k - $120k
As a Performance II-Epic, your role is to provide reliability engineering services through observability and performance engineering techniques.... ...passion for optimizing operational efficiency. You will use Site Reliability Engineering practices to deliver a seamless user...Full timePart timeWork experience placementRemote workFlexible hours$120k - $175k
...of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's possible....Full timeRemote workWork visaFlexible hours- ...itD is seeking a Site Reliability Engineer to develop and enhance automation solutions that improve the reliability, scalability, and operational efficiency of large-scale cloud infrastructure. The ideal candidate will bring hands-on experience in site reliability engineering...Work experience placementRemote work
- ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of... ...and position yourself among the top echelon in site reliability. As an Associate Site Reliability Engineer at JPMorgan Chase...Worldwide
- ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing... .... Position Summary: The Senior Azure Site Reliability Engineer acts as an advanced senior individual...Work at officeShift workDay shift
- ...is a recognized, award-winning leader in supply chain AI and a FedRAMP® authorized provider to the federal government. Site Reliability Engineer Location: U.S. (Hybrid) This role requires U.S. citizenship and eligibility for a U.S. security clearance. Role Summary...Work at officeWork from homeFlexible hours
$119k - $170k
...the greater good, come make your next move with Zscaler. Our Engineering team built the world’s largest cloud security platform from... ...cloud-first strategy. We’re looking for an experienced Staff Site Reliability Engineer (Federal) to join our Government Cloud team....Full timeWork at officeLocal areaWorldwideNight shift- ...Site Reliability Engineer Location: Schaumburg, IL or Secaucus, NJ (Hybrid) Mandatory Skills: Python/R and ML libraries (scikit-learn, TensorFlow, PyTorch), Data analysis and visualization (Pandas, NumPy, Power BI/Tableau), SQL and database management Key Responsibilities...Local area
- ...Site Reliability Engineer As a Site Reliability Engineer, your role is to provide reliability engineering services through observability and performance engineering techniques. Using monitoring and performance tools to deliver detailed feedback to product owners and...Work experience placement
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer...
- ...contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology Team, you draw upon your advanced knowledge to identify new opportunities...
- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Cloud Foundational Services team, you hold a leadership role in your team, demonstrate...
$150k - $200k
...our CEO's funding announcement: . The Reliability team owns the availability, performance,... ...enforcing reliability standards across engineering Designing incident response processes and... ...strong ownership of production systems. As a Site Reliability Engineer on the Reliability...Remote workVisa sponsorshipWork visaFlexible hours- ...meaningful products that make a real impact on children's education and literacy. About the Role We're looking for a Senior Site Reliability Engineer to drive the stability, observability, and reliability of Epic's platform as we grow. You are an experienced engineer who...Remote work
- ...are looking for people just like you. Join our team and help us develop game-changing, high-quality solutions. As a Lead Site Reliability Engineer at JPMorganChase within the Corporate sector, Enterprise Technology team, you are an integral part of a team that develops...Work at office
- ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of... ...and position yourself among the top echelon in site reliability. As a Sr Lead Site Reliability Engineer at JPMorgan Chase within...
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Chief Data & Analytics...Work at office
- ...we serve.The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted... ...scalability, and performance of enterprise platforms.As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across...Remote workFlexible hours
- UiPath, Inc. is seeking a Senior Software Engineer for our Site Reliability Engineering organization. You will design, build, and operate SRE platform systems, leveraging AI to improve reliability and performance across critical services. You will work on live-site monitoring...
$90.3k - $189.6k
...Job Title: Release Train Engineer Job Category: Information Technology Time Type: Full time Minimum Clearance Required to Start: Secret Employee Type: Regular Percentage of Travel Required: Up to 10% Type of Travel: Local The Opportunity: As the...Full timeContract workWork experience placementLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Ridgewood, NY
- site reliability engineer remote
- site reliability engineer sre
- site reliability engineering manager
- site reliability engineer
- lead site reliability engineer
- junior site reliability engineer
- site activation specialist
- website development
- on-site clinical research associate (traveling/remote)


