Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

2T Consulting

We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.

The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.

Required Technical Skills
  • Strong understanding of Site Reliability Engineering principles and operational excellence.
  • Experience with infrastructure reliability, service availability, resiliency, and performance optimization.
  • Storage Space Direct and failover clustering technical expertise. (Storage Spaces Direct enables you to build highly available, software-defined storage by pooling local disks (SSDs, NVMe drives, and HDDs) across multiple Windows Server nodes in a cluster. Instead of relying on an external SAN, S2D uses the servers' local storage to create a resilient shared storage pool)
  • Experience managing production-critical infrastructure environments with high availability requirements.
  • Experience with incident management, problem management, RCA, and continuous operational improvement.
  • Knowledge of monitoring, observability, alerting, and performance management.
Microsoft Hyper-V (Core Expertise)
  • Deep hands-on expertise in Microsoft Hyper-V architecture, deployment, administration, troubleshooting, and optimization.
  • Extensive experience in operating enterprise private cloud environments on Hyper-V.
  • Strong experience supporting enterprise-scale VDI deployments on Hyper-V.
  • Hyper-V Failover Clustering and high-availability architecture.
  • Storage integration including SAN, NAS, Storage Spaces Direct (S2D), Cluster Shared Volumes (CSV), and storage optimization.
  • Networking within Hyper-V environments including virtual switches, VLANs, NIC Teaming, QoS, and network performance tuning.
  • System Center Virtual Machine Manager (SCVMM).
Automation & Platform Engineering
  • Strong PowerShell scripting and automation experience.
  • Experience automating infrastructure deployment, operational tasks, health checks, and reporting.
  • Familiarity with Infrastructure as Code concepts and configuration management.
  • Experience developing reusable operational tooling to improve reliability and reduce manual effort.
Preferred Skills
  • Windows Server 2016/2019/2022 administration.
  • Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
  • Exposure to hybrid cloud and private cloud platforms.
  • Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
  • Experience supporting enterprise VDI environments.
  • Understanding of ITIL Incident, Problem, Change, and Release Management.
  • Experience working in regulated industries such as Banking or Financial Services.
Experience & Qualifications
  • 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
  • Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
  • Proven experience implementing automation to reduce operational overhead and improve service reliability.
  • Experience supporting enterprise private cloud and VDI environments.
  • Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
  • Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
  • Experience in Banking or Financial Services environments is advantageous.
Key Responsibilities
  • Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
  • Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
  • Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
  • Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
  • Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
  • Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
  • Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in New York, NY vacancy
  • $141k - $216.6k

     ...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical... 
    Suggested
    Work experience placement
    Work at office

    Axon

    New York, NY
    4 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Suggested
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    1 day ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Suggested
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    New York, NY
    5 days ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    3 days ago
  • $45 - $85 per hour

    DescriptionThe Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably.We're looking for engineers to be part of an empowered, self-organizing group, with the opportunity to use modern languages and tools and to operate software... 
    Suggested
    Contract work
    Temporary work

    TEKsystems

    New York, NY
    1 day ago
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating... 
    Full time

    Ridgeline

    New York, NY
    1 day ago
  • $207k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives. Mentor other engineers and contribute to the engineering community through documentation... 
    Full time
    Work at office

    Google

    New York, NY
    2 days ago
  • $110k - $120k

     ...largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading financial services and healthcare technology company based on... 
    Ongoing contract
    Full time
    Casual work
    Remote work
    Flexible hours

    SS&C Technologies

    New York, NY
    3 days ago
  • $200k - $250k

    Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the... 
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    2 days ago
  • $120k - $200k

     ...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery...  ...practices (e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)
    Overseas

    Comrise

    New York, NY
    3 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad... 
    Shift work

    JP Morgan Chase

    New York, NY
    2 days ago
  •  ...human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You... 
    Full time
    Work at office

    Dune Security

    New York, NY
    4 days ago
  • $125k - $130k

     ...Astronomer Customer Reliability Engineering RoleAstronomer empowers data teams to bring mission-critical software, analytics, and AI to life and...  ...troubleshooting skillsBonus Points If You HaveExperience as a Site Reliability EngineerWorked with Kubernetes Custom... 
    Remote work
    Weekend work

    ASTRONOMER INC

    New York, NY
    1 day ago
  •  ...Overview We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise... 

    Longfinch Technologies

    New York, NY
    3 days ago
  • $160k - $180k

     ...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience... 

    Socure Inc

    New York, NY
    1 day ago
  • $179k - $226k

     ...the team the company's Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure footprint:...  ...organization. Our challenge isn't just scale—it’s making that scale reliable, secure, and operable with less manual work. We're looking for... 
    Work at office
    Local area
    Immediate start
    Work from home
    Home office
    Monday to Friday
    Flexible hours

    United States Digital Space LLC

    New York, NY
    3 days ago
  • Gusto is hiring for a hands-on operations role focused on running and sharpening day-to-day reconciliation and loss detection. You will own execution, resolve issues, and automate manual parts using AI and data tooling. You’ll gain end-to-end understanding of money flow...

    Triwill Group

    New York, NY
    2 days ago
  •  ...Federal Reserve Bank of New York is seeking an experienced Cloud AWS Support Reliability Engineer (SRE) to build and maintain scalable AWS infrastructure and CI/CD pipelines. The role emphasizes observability, security, and resilience across enterprise cloud platforms... 

    Federal Reserve Bank (NY)

    New York, NY
    2 days ago
  •  ...subscriptions at scale, combining the agility of a high-growth business with the backing of a global organization. As the Site Reliability Engineer, you will help ensure the reliability, scalability, and observability of CloudBlue’s multi-tenant SaaS platforms used by... 
    Remote work
    Worldwide
    Flexible hours

    HostPapa

    New York, NY
    4 days ago
  •  ...The Office of Technology and Innovation (OTI) in Brooklyn, NY, seeks a Principal Automation Engineer to provide Site Reliability engineering for automation services and lead DevOps initiatives. You will write/update infrastructure code, manage CI/CD pipelines, and be... 
    Work at office

    TECHNOLOGY & INNOVATION

    New York, NY
    2 days ago
  • $150k - $170k

     ...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud‑native software applications that serve millions of customers and process billions of dollars in payments. We’re looking for... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    ZIP

    New York, NY
    5 days ago
  • $180.5k - $236.91k

     ...Senior Software Engineer, Cloud Infrastructure / SRENew York, New York, United StatesHi, we're Oscar. We're hiring a Senior Software...  ...your team's business and technical domains such as DevOps, site reliability, and cloud best practicesLead the planning, execution and release... 
    Full time
    Work at office
    Flexible hours

    Oscar Health

    New York, NY
    2 days ago
  •  ...scale our Equities business and launch Options. You will drive reliability, observability, and automation across on-prem production...  ...Development, Network, Ops, and Platform teams. You will mentor engineers, define standards, and remain hands-on to tackle complex production... 

    IEX

    New York, NY
    2 days ago
  • $115k - $160k

     ...Senior Site Reliability Engineer - AVP - Credit Trade FloorEmbark on a transformative journey as a Senior Site Reliability Engineer - AVP - Credit Trade Floor. At Barclays, our vision is clear – to redefine the future of banking and help craft innovative solutions. You... 
    Hourly pay
    Work at office

    Barclays

    New York, NY
    2 days ago
  • Altera Digital Health in New Jersey seeks a skilled IT Specialist to troubleshoot and resolve highly complex product issues, focusing on backend infrastructure and database performance. You will diagnose problems by connecting to client environments, communicate progress...

    MediSolution

    New York, NY
    5 days ago
  •  ...A financial technology company based in New York is seeking a Product & Platform Monitoring Manager to ensure the reliability and health of their fintech products. The role involves monitoring customer journeys, APIs, and incident management, requiring 5+ years of experience... 

    ClarityPay Program Services, LLC

    New York, NY
    1 day ago
  • $150k - $220k

     ...teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative...  ...can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping... 
    Local area

    Forge Global

    New York, NY
    5 days ago
  • $182k - $250.8k

     ...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    New York, NY
    1 day ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    1 day ago
  • $150k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people...  ...markets.Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience,... 
    Full time
    Temporary work
    Part time

    Goldman Sachs

    New York, NY
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!