Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

HPC Data Center Operational Lead

P2P

HPC Infrastructure Operations Lead Location: Chicago or New York (On‑site 5 days/week; regular travel to HPC data center sites required) Jump's HPC infrastructure powers some of the most demanding computational workloads in the industry. As our HPC footprint grows, we need a seasoned operations leader to own the reliability, standards, and day‑to‑day excellence of these environments. What You'll Do: Team Leadership & Organizational Ownership Lead and manage data center site leads and their teams across multiple HPC facilities; site leads report directly to this role. Recruit, mentor, and develop team members while conducting performance reviews and building a culture of operational rigor. Direct onsite contractors by providing clear scope and validating completed work. HPC Data Center Standards, Processes & Preventative Maintenance Develop, document, and enforce operational standards and procedures for Jump's HPC data centers covering power, cooling, cabling, and hardware lifecycle. Design and own the preventative maintenance program, including scheduled inspections, component replacements, and firmware/capacity reviews to minimize unplanned downtime. Drive continuous improvement of operational processes and pursue automation—including AI‑driven approaches—to reduce manual effort and human error. Critical Facility Systems Expertise Serve as the subject matter authority on HPC data center power distribution, power striping strategies, and failover/redundancy configurations. Own expertise across air cooling, liquid cooling (direct‑to‑chip, rear‑door, CDU‑based), and hybrid cooling architectures. Maintain deep knowledge of environmental monitoring and controls (temperature, humidity, airflow, leak detection) and ensure systems remain within design parameters. Monitoring & Incident Response Own the HPC data center monitoring strategy end‑to‑end: define what is monitored, set alerting thresholds, and ensure comprehensive visibility into facility and hardware health. Leverage AI tools to analyze telemetry data, identify failure patterns, predict potential issues, and accelerate root cause analysis during incidents. Lead critical incident response and drive root cause analysis and corrective actions to prevent recurrence. Establish and track operational KPIs including availability, mean time to repair, and efficiency metrics. Server & Switch Hardware Expertise Maintain deep, hands‑on knowledge of server hardware architectures including multi‑socket platforms, GPU/accelerator configurations, memory subsystems, NVMe/storage controllers, BMC/IPMI management, and firmware lifecycle. Maintain deep, hands‑on knowledge of network switch hardware including line cards, optics/transceivers, switch fabrics, and platform‑specific diagnostics for Arista and Cisco platforms. Evaluate new hardware platforms, drive hardware qualification and acceptance testing, and provide informed recommendations on hardware selection. Hardware Break‑Fix Own the overall hardware break‑fix function across all HPC sites, ensuring rapid diagnosis and resolution for servers, GPUs, network equipment, storage, and facility infrastructure. Diagnose complex hardware failures at the component level—CPUs, DIMMs, GPUs, NICs, PSUs, fans, drives, switch line cards, and optics—and direct the team to resolve efficiently. Establish escalation paths, SLA targets, and reporting for hardware failures. Inventory & Spares Management Own inventory processes and spares tracking across all HPC facilities, ensuring critical spares are stocked, tracked, and replenished to meet availability targets. Maintain accurate asset records for all serialized and consumable inventory. Planning, Vendor & Budget Management Conduct capacity planning for space, power, cooling, and cabling to stay ahead of growth. Gather requirements and plan new hardware installations including physical placement, power/cooling needs, and cabling. Manage relationships with colocation providers and hardware vendors; negotiate contracts and SLAs. Develop and manage operational budgets for equipment, staffing, and facilities. Networking & Linux Possess strong working knowledge of networking concepts including L2/L3 protocols, VLANs, BGP, OSPF, LACP, ECMP, and high‑performance fabrics relevant to HPC environments. Understand network architectures such as spine‑leaf, fat‑tree, and high‑radix topologies used in HPC clusters. Maintain strong Linux systems knowledge—comfortable navigating and troubleshooting at the OS level, including storage, networking, process management, log analysis, and system diagnostics. AI‑Driven Operations Use AI tools daily across all aspects of the role: writing and reviewing documentation, analyzing operational data, drafting procedures, managing communications, and problem‑solving. Champion AI adoption within the team—set the expectation that every team member integrates AI into their daily workflows. Identify and implement opportunities where AI can replace or augment manual operational processes. Cross‑Team Partnership Partner with HPC Engineering, Network Engineering, and other teams to align operations with research and business needs. Ensure compliance with all safety, security, and regulatory requirements. Travel Travel regularly to Jump's HPC data center sites for operational oversight, project execution, and team engagement. This is a core requirement of the role. Additional duties as assigned or needed. Skills You'll Need: Minimum 7+ years of data center operations experience with at least 3 years leading teams in 24/7 critical infrastructure environments. HPC environment experience strongly preferred. In‑depth knowledge of data center power systems, power distribution/striping, and failover/redundancy architectures. In‑depth knowledge of cooling technologies including air cooling, liquid cooling (direct‑to‑chip, rear‑door heat exchangers, CDUs), and environmental control systems. Proven experience building and maintaining preventative maintenance programs and operational standards/procedures. Strong experience with data center monitoring platforms (DCIM, BMS, environmental sensors) and defining monitoring/alerting strategies. Demonstrates a high level of energy, results driven, and able to work under pressure with tight deadlines. Technical Skills: Deep knowledge of server hardware architectures: multi‑socket platforms, GPU/accelerator systems, memory subsystems, NVMe storage, BMC/IPMI, and firmware management. Deep knowledge of network switch hardware: line cards, optics/transceivers, switch fabrics, and platform diagnostics across Arista and Cisco platforms. Proven hardware break‑fix experience with the ability to diagnose failures at the component level (CPUs, DIMMs, GPUs, NICs, PSUs, drives, line cards, optics). Strong understanding of networking concepts: L2/L3 protocols, VLANs, BGP, OSPF, LACP, ECMP, and HPC network topologies (spine‑leaf, fat‑tree). Strong Linux systems proficiency—well beyond basic CLI usage. Comfortable with OS‑level troubleshooting, storage and network configuration, process management, log analysis, and system diagnostics. Experience managing inventory and spares programs for critical infrastructure. Structured cabling standards expertise. Programming/scripting experience (Python preferred) is a plus. Demonstrated heavy use of AI tools (e.g., LLM‑based assistants, AI coding tools, AI‑driven analytics) in a professional setting. You should already be using AI daily and be eager to push its application further across operations. Strong project management skills with multi‑site infrastructure deployment experience. Knowledge of industry standards including ASHRAE and TIA‑942. Excellent written and verbal communication skills with the ability to communicate effectively across technical and non‑technical audiences. Meet physical requirements including working on ladders/elevated platforms and lifting up to 50 lbs. Extremely high personal standards for work quality and operational discipline. Reliable and predictable availability, including ability to work evenings and weekends as required. Willingness and ability to travel regularly to data center sites. Bachelor's degree preferred. #J-18808-Ljbffr P2P

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the HPC Data Center Operational Lead in New York, NY vacancy
  • P2P is seeking an HPC Infrastructure Operations Lead to manage critical infrastructure and ensure the reliability of our data centers. This on-site role in Chicago or New York requires leadership and technical expertise in data center operations, including a strong focus... 
    Suggested

    P2P

    New York, NY
    15 hours ago
  • $125k - $150k

     ...electric vehicles, renewable energy, and data centers. KoBold builds AI models for mineral...  ...—to guide decisions on KoBold-owned-and-operated exploration programs. In the six years since...  ...and software engineers, who come from leading technology companies, jointly lead... 
    Suggested
    Full time
    Contract work
    For contractors
    Seasonal work
    Local area
    Remote work

    KoBold Metals

    New York, NY
    2 days ago
  • $160k - $190k

     ...working, and collegial? Join us at SB Energy, a leading company backed by SoftBank and Ares pairing...  ...City, CA, SB Energy develops, builds, owns & operates some of the largest and most technically advanced energy and data center infrastructure projects in the United States.... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Local area
    Flexible hours

    SB Energy

    New York, NY
    2 days ago
  • $157k - $210k

     ...AI with confidence. Trusted by leading AI labs, startups, and global enterprises...  ...more at .What You'll Do:The DC (Data Center) Program and Cost Management team is the operational planning team responsible for...  ...with hyperscale or AI/HPC data center environments, including... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    2 days ago
  • $150k - $250k

     ...every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and...  ...arm of senior operations leadership in the field - leading site assessments and operational audits, driving team technical... 
    Suggested
    For contractors
    Local area
    Shift work

    Fluidstack

    New York, NY
    2 days ago
  • $250k

     ...Maryland, Virginia market Vice President, Data Center Operations Location: San Francisco, New York City,...  ...is building GPU supercomputers for leading AI labs, enterprises, and government organizations...  ...: Experience with GPU clusters, HPC environments, or hyperscale/cloud... 
    Full time

    LVI Associates

    New York, NY
    2 days ago
  • $143k - $210k

     ...and scale AI with confidence. Trusted by leading AI labs, startups, and global...  ...025. Learn more at .What You’ll Do:The Data Center Security organization at CoreWeave is responsible...  ...scalable physical security programs, operational rigor, and strong cross-functional partnership... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    For contractors
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    2 days ago
  • $86.21k - $141.64k

    The Operational Risk Program Lead is a key contributor to how the organization understands, manages, and communicates operational risk. This role...  ...including operational, third-party, technology, cyber, model, data, and AI risk by helping translate complex risk information... 
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Guardian Life Insurance

    New York, NY
    3 days ago
  • $134k - $179k

     ...and scale AI with confidence. Trusted by leading AI labs, startups, and global...  ...Program and Cost Management team is the operational planning team responsible for the program...  ...across the entire lifecycle of CoreWeave's data centers, ensuring strategic program objectives... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    15 hours ago
  • ETAP is seeking an execution-focused Project Manager to lead real-time microgrid control and power management solutions for hyperscale data center clients. You will own projects end-to-end from contract through commissioning and acceptance, acting as the accountable leader... 
    Contract work

    AVEVA Denmark

    New York, NY
    3 days ago
  • TEKsystems in New Jersey is seeking a Data Center Project Manager to lead complex MEP and critical-infrastructure capital projects across enterprise...  ..., installation, commissioning, and handover to operations while ensuring performance, safety, and regulatory compliance... 

    TEKsystems

    New York, NY
    2 days ago
  •  ...Talent Partners is seeking an experienced MEP Project Manager for data center construction projects. You will oversee planning, coordination,...  ...preconstruction through commissioning and turnover. You will lead a team of engineers and field staff, coordinate with design... 

    Project Talent Partners

    New York, NY
    2 days ago
  • $195k - $235k

    COMPANY OVERVIEW KKR is a leading global investment firm that offers alternative asset...  ...subsidiaries. TEAM OVERVIEW KKR’s Global Data Operations team is responsible for collecting,...  ...Within Data Operations, the Data Operations Center of Excellence (CoE) operates as a hub‑... 
    Local area

    Stage

    New York, NY
    3 days ago
  • SB Energy, a growing energy and data center infrastructure company, seeks a Site Operations Lead to run one or more data center campuses. You will manage a team of 3rd party CFEs and internal staff, interface with Engineering & Construction, and direct maintenance, events... 

    SB Energy

    New York, NY
    2 days ago
  •  ...to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building... 
    Contract work

    Fluidstack

    New York, NY
    1 day ago
  •  ...Management and Compliance, you are at the center of keeping JPMorgan Chase strong and...  ....As a Technology and Cybersecurity Operational Risk Management Lead with Compliance, Conduct and...  ...and Compliance Risk teams; developing data-driven approaches that leverage agentic... 

    JP Morgan Chase

    New York, NY
    1 day ago
  •  ...GardaWorld company, is widely regarded as the leading integrated risk management, crisis...  .... Championed by our advanced Global Operation Centers and our skilled team of intelligence analysts...  ...Information Security: Protect the data and systems of Crisis24 and its stakeholders... 
    Local area
    Flexible hours
    Night shift

    Onsolve by Crisis24

    New York, NY
    1 day ago
  •  ...Job Description Job Description Operations Manager, Belong Center Location: Greenpoint Brooklyn, New York Reports to: DAYBREAKER Founder...  ...each project has the resources it needs to thrive. Lead the completion of annual nonprofit compliance obligations... 

    DAYBREAKER

    New York, NY
    4 days ago
  •  ...Job Description Job Description ICE OPERATIONS SUPERVISOR- Davis Center at the Harlem Meer Sports Facilities Management, LLCLOCATION: New York...  ...Facilities Companies (SFC) company. SFC is the nation's leading resource for managing and developing sports, recreation,... 
    Seasonal work
    Night shift
    Weekend work

    The Sports Facilities Companies

    New York, NY
    8 days ago
  • $28 per hour

     ...environments into functional, beautiful spaces—and we’re looking for an Operations & Team Manager (OTM) to supervise and coach on-site staff,...  ...’ll share a key takeaway with the team, and bi-weekly, they’ll lead a brief discussion during evaluations on how to apply these... 
    Full time
    Trial period
    Monday to Friday
    Shift work

    ORGANIZED HUMAN LLC

    New York, NY
    2 days ago
  •  ...RCM Technologies, Inc. in New York City seeks a Tenant Engagement Director to build and own the tenant relationships for data-center tenants. You will design the function from the ground up, ensuring SLA rigor and financial accountability from day one. You will drive... 

    RCM Technologies, Inc.

    New York, NY
    19 hours ago
  • $175k - $225k

     ...seeking an experienced Director of Information Technology to lead and transform our IT operations as we separate the IT function from Security. This...  ...(Microsoft Azure, IaaS/PaaS), on‑premises data center resources, and enterprise networks. Implement highly resilient... 
    Contract work
    Work at office
    Flexible hours

    Pathstone

    New York, NY
    2 days ago
  • $200k - $250k

     ...Lead Project Manager - Mission Critical Lead | Data Centers, NYC - $200k - $250k About the Company:  I am partnered with a New York City based General Contractor /Construction Management firm achieving YoY growth since inception building a mix of ground-up and interiors... 
    Temporary work
    For contractors
    Work at office

    ArtemusPoint

    New York, NY
    4 days ago
  • Position: Operations Director Location: Valley Green, PA Amphenol Communications Solutions...  ...including server, storage, data center, mobile, RF, networking, industrial, business...  ...platforms, helping power the technology behind leading Tier 1 OEMs. With global design, sales,... 

    Amphenol ICC

    New York, NY
    2 days ago
  • JAC Recruitment is partnering with a well-established Japanese financial institution in New York to recruit an Assistant Vice President for its Risk Management team. The role focuses on governance, risk reporting, policy development, and committee support in a hybrid office...

    JAC Recruitment

    New York, NY
    19 hours ago
  •  ...- Enterprise Infrastructure, Network, Data Center, Cloud, Security & Compliance Experience...  ...responsible for providing strategic and operational leadership for the organization’s...  ...cybersecurity, and infrastructure operations while leading large-scale data center migrations,... 
    Remote work
    Relocation

    Infinite Computer Solutions

    New York, NY
    1 day ago
  • $110k - $150k

     ...is seeking a dedicated Architectural Project Manager for their Data Centers team in New York City. This full-time on-site role requires extensive...  ...experience. The successful candidate will manage projects, lead teams, and ensure quality control throughout project lifecycles... 
    Full time

    Corgan

    New York, NY
    2 days ago
  • CRB is seeking an Operational Excellence Manager in Georgia to enhance project delivery processes within the organization. This role focuses on driving continuous improvement initiatives and effective communication with stakeholders to document and standardize operational... 

    CRB

    New York, NY
    4 days ago
  • Johnson Matthey is seeking a Lead Operation Excellence Specialist in Devon, PA to contribute to its mission of sustainable technology. This role focuses on driving continuous improvement and operational performance, working with site management to identify priorities and... 

    Johnson Matthey

    New York, NY
    1 day ago
  • White Cap is seeking a dedicated professional to drive operational excellence and continuous improvement initiatives. This role involves...  ...and expertise in Lean principles. Key responsibilities include leading improvement initiatives, facilitating workshops, and... 

    White Cap

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to HPC Data Center Operational Lead. Be the first to apply!