Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

2T Consulting

We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.

The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.

Required Technical Skills
  • Strong understanding of Site Reliability Engineering principles and operational excellence.
  • Experience with infrastructure reliability, service availability, resiliency, and performance optimization.
  • Storage Space Direct and failover clustering technical expertise. (Storage Spaces Direct enables you to build highly available, software-defined storage by pooling local disks (SSDs, NVMe drives, and HDDs) across multiple Windows Server nodes in a cluster. Instead of relying on an external SAN, S2D uses the servers' local storage to create a resilient shared storage pool)
  • Experience managing production-critical infrastructure environments with high availability requirements.
  • Experience with incident management, problem management, RCA, and continuous operational improvement.
  • Knowledge of monitoring, observability, alerting, and performance management.
Microsoft Hyper-V (Core Expertise)
  • Deep hands-on expertise in Microsoft Hyper-V architecture, deployment, administration, troubleshooting, and optimization.
  • Extensive experience in operating enterprise private cloud environments on Hyper-V.
  • Strong experience supporting enterprise-scale VDI deployments on Hyper-V.
  • Hyper-V Failover Clustering and high-availability architecture.
  • Storage integration including SAN, NAS, Storage Spaces Direct (S2D), Cluster Shared Volumes (CSV), and storage optimization.
  • Networking within Hyper-V environments including virtual switches, VLANs, NIC Teaming, QoS, and network performance tuning.
  • System Center Virtual Machine Manager (SCVMM).
Automation & Platform Engineering
  • Strong PowerShell scripting and automation experience.
  • Experience automating infrastructure deployment, operational tasks, health checks, and reporting.
  • Familiarity with Infrastructure as Code concepts and configuration management.
  • Experience developing reusable operational tooling to improve reliability and reduce manual effort.
Preferred Skills
  • Windows Server 2016/2019/2022 administration.
  • Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
  • Exposure to hybrid cloud and private cloud platforms.
  • Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
  • Experience supporting enterprise VDI environments.
  • Understanding of ITIL Incident, Problem, Change, and Release Management.
  • Experience working in regulated industries such as Banking or Financial Services.
Experience & Qualifications
  • 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
  • Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
  • Proven experience implementing automation to reduce operational overhead and improve service reliability.
  • Experience supporting enterprise private cloud and VDI environments.
  • Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
  • Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
  • Experience in Banking or Financial Services environments is advantageous.
Key Responsibilities
  • Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
  • Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
  • Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
  • Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
  • Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
  • Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
  • Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in New York, NY vacancy
  • $140k - $170k

     ...data security, navigate complex regulatory compliance and optimize business interactions. Role Description: As a Site Reliability Engineer, you will work with Agile engineering teams to provide production insight into running and operating software at-scale in... 
    Suggested
    Full time
    Local area

    Symphony Communication Services

    New York, NY
    16 hours ago
  • $123k - $165k

    Job Summary:Department/Group OverviewOur engineering fleet is a horizontal set of teams...  ...organization. Our specific team provides reliability engineering and operational support to backend...  ...products and brands.We are seeking a Site Reliability Engineer who will contribute... 
    Suggested

    Disney Interactive

    New York, NY
    5 days ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    20 hours ago
  • $165k - $241.4k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Suggested
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    4 days ago
  • $141k - $216.6k

     ...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical... 
    Suggested
    Work experience placement
    Work at office

    Axon

    New York, NY
    3 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Full time
    Work at office
    3 days per week

    GrubHub

    New York, NY
    4 days ago
  • $131k - $164k

    Position Overview We are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation frameworks, to join our global Infrastructure & Operations team. This role is a hands-on senior engineering position responsible... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    Diligent Corporation

    New York, NY
    2 days ago
  • $120k - $200k

     ...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery...  ...practices (e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Overseas

    Comrise

    New York, NY
    2 days ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    20 hours ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    New York, NY
    5 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad... 
    Shift work

    JP Morgan Chase

    New York, NY
    20 hours ago
  • $138.1k - $198.2k

     ...more intuitive with technology that simply works.  The SRE Engineering Enablement Team supports our CI Platforms, Developer Environments...  ...Our customers are all engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter of our engineering... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    New York, NY
    4 days ago
  • $110k - $120k

     ...largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading financial services and healthcare technology company based on... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    20 hours ago
  • $104.9k - $174.7k

     ...Management. You can learn more about LexisNexis Risk at the link below, About the Role: We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    LexisNexis Risk Solutions

    New York, NY
    4 days ago
  •  ...human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You... 
    Full time
    Work at office

    Dune Security

    New York, NY
    3 days ago
  •  ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client... 

    Karsun Solutions

    New York, NY
    5 days ago
  •  ...researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience. The Role: At...  ...new cloud infrastructure company, we seek to improve our reliability dramatically while scaling the size of our platform and customer... 

    Modal Labs

    New York, NY
    5 days ago
  •  ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1... 

    Karsun Solutions

    New York, NY
    5 days ago
  •  ...A dynamic fintech company in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability of its fintech products. The role focuses on end-to-end monitoring of customer journeys, incident management, and collaboration with various teams. Candidates... 

    ClarityPay

    New York, NY
    5 days ago
  •  ...No one coasts. If you're driven by impact, pace, and raising the bar. This is the place. The role As a Senior Site Reliability Engineer you'll join the founding SRE team at our new NYC engineering hub, sitting within Foundations. You'll own critical services... 
    Work at office

    Legora

    New York, NY
    4 days ago
  • $100k - $250k

     ...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial...  ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics... 
    Local area

    Kalshi Inc

    New York, NY
    4 days ago
  • $225k - $325k

     ...What you'll do day-to-day Ensure the scalability, reliability, and observability of our systems to maintain and improve the firm's core infrastructure environment. Lead a range of engineering projects, from developing proprietary platforms for configuration... 
    Hourly pay

    D. E. Shaw & Co.

    New York, NY
    2 days ago
  • $155k - $222.6k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications ~2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    New York, NY
    1 day ago
  •  ...Asia and the Middle East. We are creative, low-ego and team-spirited. The Role We are seeking highly experienced Site Reliability Engineers (SRE) to shape the reliability, scalability and performance of our platform and customer facing applications. You will work... 
    Relocation package

    Mistral Ai

    New York, NY
    3 days ago
  •  ...Job Title WHAT YOU'LL DO DAY-TO-DAY: The engineer will be responsible for developing and implementing new systems and services in the areas of infrastructure monitoring, configuration management, and automation. Additional responsibilities include upgrading and/or... 

    1872 Consulting

    New York, NY
    4 days ago
  • $160k - $180k

     ...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience... 

    Socure Inc

    New York, NY
    5 days ago
  • $111k - $218k

     ...The Site Reliability Engineering team designs and builds the global infrastructure on which we deploy our services, focusing on the above mentioned flagship MongoDB Atlas platform. As our customers grow and globalize, our services must satisfy demands for low-latency... 
    Local area
    Worldwide
    Flexible hours

    MongoDB

    New York, NY
    4 days ago
  •  ...Komodor, a remote-first company, is seeking a Solutions Engineer to connect customer business initiatives to the Komodor platform, understanding developers, DevOps and Incident response teams working with Kubernetes. You will identify customer pain points and communicate... 
    Remote work

    Komodor

    New York, NY
    5 days ago
  •  ...Participate in an oncall rotation. Work with teams across the company to ensure we achieve the right balance of developer velocity, reliability and performance, and cost efficiency. What You'll Bring ~5+ years of experience ~ Experience with containerization and... 

    Clay Labs

    New York, NY
    4 days ago
  • $115k - $125k

     ...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help... 
    Work experience placement
    Work at office
    Work from home
    Monday to Friday
    Flexible hours
    Shift work
    Weekend work
    Afternoon shift
    Early shift

    Pico

    New York, NY
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!