Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

$179k - $218k
Full-time

Crusoe

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

The Mission

Crusoe is building the world’s most climate-aligned AI infrastructure. As we scale toward unprecedented power densities and liquid-cooled architectures, the gap between "Data Center Design" and "Silicon Reality" must be bridged.

We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority on GPU platforms within the Data Center Engineering and Operations organization. Your mission is twofold: act as the primary technical consultant to our Data Center Engineering team to ensure future facilities are built for next-gen silicon, and provide the Operations team with the specialized tooling, SOPs, and predictive strategies needed to maintain peak cluster health.

The Strategic Bridge

  • For DC Engineering: You are the internal consultant. You translate upcoming GPU power/thermal roadmaps (NVIDIA/AMD) into design requirements for our next-generation facilities.
  • For Site Operations: You are the "Technical Enabler." You develop the diagnostic tools and technical SOPs that enable field technicians to resolve complex GPU issues with surgical accuracy.
  • For Sourcing: You are the "Technical Strategist." You define the technical sparing requirements and site-level inventory needs based on hardware failure telemetry.

Key Responsibilities

  • Engineering Education & Design Support: Provide deep-dive technical guidance to the Data Center Engineering team on upcoming silicon (e.g., NVIDIA Blackwell/Rubin, AMD MI350/400). Ensure future facility designs for power, cooling, and rack-spacing are ready for 2000W+ per-chip densities.
  • Predictive Operations & Telemetry: Leverage AI/ML methodologies to analyze fleet-wide telemetry (power draws, thermal gradients, and error rates). You will lead the transition from reactive troubleshooting to predictive maintenance , identifying "pre-failure" patterns in HBM or NVLink components before they impact customer training runs.
  • Technical Sparing Architecture: Architect the site-level sparing strategy from a technical perspective. Use failure telemetry and MTBF data to define the "Critical Spares List" and stocking levels required at each site to meet cluster uptime targets, providing these requirements to Sourcing for execution.
  • Operational Tooling & SOPs: Build the "Operational Blueprint" for the field. Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or thermal throttling.
  • Advanced Troubleshooting & RCA: Act as the Tier-3 escalation point for the most complex hardware failures in the production environment. Lead Root Cause Analysis (RCA) on systemic issues that span the boundary between hardware and facility environmental factors.
  • Silicon Roadmap Authority: Maintain a 24-month forward-looking view of NVIDIA and AMD architectures. Educate internal stakeholders on how transitions in HBM4, interconnect speeds, and liquid-cooling will impact Crusoe’s physical infrastructure.
  • Vendor & VAR Technical Lead: Support the technical relationship with OEMs and VARs. Audit their hardware builds, review their technical bulletins, and ensure their hardware roadmaps align with Crusoe’s operational and engineering standards.

Technical Requirements

  • Silicon & Fabric Mastery: Expert-level knowledge of NVIDIA (Hopper/Blackwell/Rubin) and AMD (Instinct) architectures. Mastery of the physical and logical layers of NVLink, NVSwitch, and InfiniBand.
  • Infrastructure Bridge-Building: Ability to translate "Silicon Data Sheets" into "Mechanical Engineering Requirements." You can explain how a GPU's specific heat-load profile affects CDU sizing and secondary loop design.
  • Data-Driven Diagnostics: Proficient in Python, Go, or Bash to build telemetry and health-check tools (utilizing DCGM and ROCm ). Experience using large datasets or basic ML frameworks to build "Smart Monitoring" that filters critical health signals from noise.
  • Operational Reliability Analysis: Experience using failure telemetry to inform site-level sparing requirements and field-service workflows.
  • Thermal Management: Deep understanding of the operational realities of Direct-to-Chip (D2C) cooling, including fluid dynamics, pressure-drop curves, and the lifecycle of dripless couplings.

Qualifications

  • 10+ years in Hardware Engineering, Systems Architecture, or Data Center Infrastructure.
  • The "Consultant" Mindset: Proven track record of educating and influencing cross-functional teams (specifically Engineering and Operations).
  • GPU Authority: You have managed or architected GPU clusters at scale (thousands of nodes) at a hyperscaler, a GPU-specialized cloud, or a major silicon vendor.
Education: B.S. or M.S. in Electrical Engineering, Computer Engineering, or a related technical field. Benefits:
  • Competitive compensation
  • Restricted Stock Units
  • Paid time off & paid holidays
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off
Compensation Range

Compensation will be paid in the range of up to $179,000 -$218,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data. (#INDDIG)

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Senior Staff Data Center Operations Engineer, GPU Hardware Architecture in Remote vacancy
  • $243.29k - $295.25k

     ...Infrastructure Foundation Hardware Engineering team, you will play a key role...  ...the technical lead for our GPU and AI accelerator...  ...GPU hardware, from initial architectural evaluation and firmware qualification...  ...understanding of modern data center technologies, including PCIe... 
    Senior
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    a month ago
  •  ...experiences—from AI and data centers, to PCs, gaming and...  ...'s Data Center GPU organization is transforming...  ...to join the Product Architecture and Workload Strategy...  ..., and performance engineering teams to translate workload...  ...collaborating across hardware, software, and system... 
    Suggested
    Remote work

    AMD

    San Jose, CA
    a month ago
  •  ...seeking a Principal Hardware Systems...  ...Platforms to provide senior technical...  ...for the architecture and development...  ...high-performance data center networking platforms...  ...and engineering leaders to define...  ...enterprise IT/network-operations role. The...  .... ~ GPU scale-up fabrics... 
    Suggested
    Local area
    Remote work

    Celestica International LP

    Remote
    25 days ago
  •  ...developers and enterprises from data and model training...  .... Built by engineers, for engineers. From large-scale GPU orchestration to inference...  ...deep expertise across hardware, software and AI R&D....  ...that power how Nebius operates its data center infrastructure at scale... 
    Senior
    Remote job
    Full time

    Nebius

    Remote
    9 days ago
  • $158k - $219.45k

     ...support for hospitals, data centers, remote sites,...  ...modern software engineering to rapidly...  ...Radiant is seeking a Senior Thermofluids Systems...  ...power conversion architecture. This...  ...You will define operating points that balance...  ...engineers owning hardware and subsystems to... 
    Senior
    Full time
    Summer work
    Immediate start
    Remote work
    Flexible hours
    Weekend work

    Radiant Industries

    El Segundo, CA
    11 days ago
  • $284.9k - $427.3k

     ...Technologies, Inc.## **Job Area:**Engineering Group, Engineering Group >...  ...*General Summary:**Qualcomm Data Center team is developing High...  ...understanding of CPU Server architecture and have a passion for architecting...  ...field and 10+ years of Hardware Engineering, Software... 
    Senior
    Work experience placement
    Work from home

    Jobleads-US

    Santa Clara, CA
    2 days ago
  •  ...search of a highly technical Senior Product Manager to lead the...  ...infrastructure, compute, and GPU platforms. This role is...  ...strategy. You’ll work closely with engineering, architecture, sales, customers, and...  ...platforms, virtualization, GPUs, or hardware-focused products. *... 
    Senior
    Remote work

    Nomad Search Partners

    Florida, FL
    7 days ago
  • $255k - $340k

     ...One person, one GPU.If you'd like to...  ...currently Tuesday.Architecture leads the design,...  ...that maximizes the data center power footprint,...  ...re looking for a Senior HPC Systems Architect...  ...and mentoring to engineering teams, fostering...  ...knowledge of HPC hardware including GPU... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  • $152k - $241.5k

     ...NVIDIA's invention of the GPU 1999 sparked the growth of...  ...company”.We are searching for a Senior Backend Compiler Engineer with experience in LLVM...  ...complex parallel SIMT architectures.What you will be doing:Guide...  ...with global compiler, hardware and application teams to oversee... 
    Senior
    Full time
    Remote work
    Worldwide

    Nvidia

    Austin, TX
    a month ago
  •  ...transport logistics, data center design and construction...  ..., and daily operations. Bitdeer also offers...  ...building an AI-operated GPU cloud spanning 4 US DCs...  ...VxLAN across spine-leaf architectures. IP transit, peering...  ...data center network engineering, with hands-on EVPN-... 
    Senior
    Full time
    Local area

    Bitdeer Technologies Group

    Remote
    9 days ago
  •  ...larger than GPUs. This architecture allows Cerebras to...  ...times faster than GPU-based hyperscale...  ...for delivering data center capacity from forecast...  ...go-live.You will operate as the SSOT (...  ...network topology, and hardware rollout schedules....  ...signal updates to senior executives.Operate... 
    Senior
    Contract work
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    a month ago
  • $255k - $340k

     ...superintelligence. One person, one GPU.If you'd like to...  ...is currently Tuesday.Hardware Engineering at Lambda is...  ...lifecycle: roadmap and architecture, proof-of-concept for...  ...in the industry operate at, this is that team...  ...deployment for HPC, data center, or cloud infrastructure... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  • $300 per month

     ...up, we own and operate each layer of the...  ...manufacturing, data center construction, and...  ...We are seeking Senior Software Engineers to design and develop...  ...bring server hardware, switches, and...  ...contributing to architecture and design (patterns...  ...-on exposure to GPU clusters or high... 
    Senior
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    20 hours ago
  • $255k - $340k

     ...One person, one GPU.If you'd like to...  ...standards and reference architectures for mechanical...  ...large-scale AI data centers.Develop control narratives...  ..., sequences of operation, points lists,...  ..., network, hardware, construction, commissioning...  ...effectively with engineers, programmers,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 days ago
  •  ...With 33 global cloud data center locations, Vultr is trusted...  ...Cloud Compute, Cloud GPU, Bare Metal, and...  ...experienced Storage Operations Engineer to build, maintain, and...  ...server-class storage hardware, networking...  ...participate in the design and architecture decisions for Ceph-... 
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Vultr

    Remote
    11 days ago
  •  ...datacenters. Our differentiated architecture seamlessly integrates hardware, software and system...  ...the efficiency of GPU, CPU and accelerator-based...  ...thinking team of architects, engineers, and business professionals...  ...within Manufacturing Operations. Supporting activities include... 
    Senior
    Full time
    Contract work
    Work experience placement
    Remote work
    Flexible hours

    Cornelis Networks

    Wayne, PA
    3 days ago
  •  ...Staffing is seeking a GPU Server Test Software Engineer to join a client...  ...Jenkins. Operating Systems: Strong...  .... Hardware & Networking: Fundamental...  ...server and network architectures with proficiency in...  ...Experience in server, data center, or test system... 
    Full time
    Local area
    Remote work

    Ultimate Staffing

    Fort Worth, TX
    3 days ago
  •  ...GPU Server Test Software Engineer Reports to: Manager – Testing Tech...  ...working with servers, data centers, or test system...  ..., and Jenkins. • Operating Systems: Proficient...  ...environments. • Hardware/Networking: Fundamental...  ...of server/network architecture and skill in remote... 
    Full time
    Casual work
    Work at office
    Remote work
    Flexible hours

    Donato Technologies Inc

    Fort Worth, TX
    2 days ago
  • $131k - $175k

     ...industry leader in data-driven,...  ...large data center, campus and...  ...such as Best Engineering Team, Best...  ...Work With As a Senior Rack...  ..., including hardware, software, thermal...  ...generation rack architectures optimized...  ...-density GPU environments...  ...briefs) used by operators and... 
    Senior
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    a month ago
  • $243.29k - $295.25k

     ...Infrastructure Foundation Hardware Engineering team, you will help develop...  ...hardware engineering, system architecture, or infrastructure...  ...server environments, modern data center technologies, and server architecture...  ...bus-level traces/captures.GPU Architectures: Familiarity... 
    Senior
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    a month ago
  • $133.5k - $184.8k

     ...asset support for hospitals, data centers, remote sites, and...  ...leverages modern software engineering to rapidly deliver safe, factory...  ...Role We are seeking a Senior Reactor Operations Engineer to design and implement...  ...team to translate complex hardware and software designs into... 
    Senior
    Full time
    Summer work
    Immediate start
    Remote work
    Flexible hours
    Weekend work

    Radiant Industries

    El Segundo, CA
    21 days ago
  • $153k - $242k

     ...About the Role As a Senior Software Engineer within our Compute Architecture organization, you will...  ...software control plane for hardware lifecycle management across large-scale GPU data centers. The METALDEV team...  ...health, automate safe operational workflows, and give operators... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    18 days ago
  •  ...businesses consume data and we...  ...breakthroughs in efficient operations. With our...  ...defines and engineers high-speed networking...  ...scale-out architectures for optimizing...  ...configurations. ~ AI Hardware Architecture:...  ...with GPU/accelerator clusters and data center infrastructure... 
    Senior
    Temporary work
    Remote work
    Flexible hours
    Shift work

    Sandisk

    Milpitas, CA
    17 days ago
  • $80k - $100k

     ...installation diagrams for data center automation and...  ...support of our solutions architecture and sales team....  ...assisting the central engineering team in the design and...  ...Designer works under the Hardware Engineering Manager,...  ...design and operation to customers, consultants... 
    Temporary work
    For subcontractor
    Work at office
    Local area
    Remote work

    Albireo Energy

    Gambrills, MD
    more than 2 months ago
  •  ...bottlenecks and driving architectural improvements?Eager to...  ...performance analysis engineers to join our diverse team...  ..., and propose data-driven architectural or...  ...analysis and debug using hardware emulation platforms.Hands...  ..., including CPU, GPU, and I/O workloads.Deep... 
    Work at office
    Local area
    Remote work

    ARM

    Austin, TX
    a month ago
  •  ...Job Title : Data Center Operations Engineer Citizenship: US Citizen or Green Card candidates only...  ...• RoCE network experience • Nvidia GPU/NIC experience is a plus New Scope...  ...resolutions • Support firmware upgrades and hardware break/fix activities • Handle parts... 
    Remote work
    Flexible hours
    Weekend work

    AMroute LLC

    Memphis, TN
    4 days ago
  • $272k - $431.25k

     ...'s first Orbital Data Center (ODC) module — a...  ...compute platform engineered for low-Earth orbit...  ...system software architecture for Space-1 and successor...  ...the host OS, GPU and CPU drivers,...  ...platform that operates reliably in the radiation...  ...with the orbital hardware system... 
    Full time
    Work experience placement
    Remote work

    NVIDIA

    Santa Clara, CA
    20 hours ago
  •  ...Description Job Description Role name: Architect/Senior Electronics Hardware Design Engineer Work Location: United states of America (remote)...  ...for embedded hardware platforms. Develop hardware architectures, schematics, interface circuits, and subsystem designs... 
    Senior
    Remote job
    Immediate start

    PDDN INC.

    Dallas, TX
    29 days ago
  • $91.4k - $187k

     ...professional with strong technical, operational, and customer support...  ...global teams across multiple seniority levels to improve functions, projects, and data center operations. Ability to travel...  ...between project teams, data center engineering, and operations to manage... 
    Senior
    Full time
    Remote work
    Flexible hours

    Oracle

    Chicago, IL
    5 days ago
  • $86.8k - $165.2k

     ...leading businesses, world-class operations and investments in research...  ...of experience and renowned engineering expertise to meet the needs...  .../Exciter & Processing Architecture (REPA) Department is the focal...  ...is for a multi-discipline Senior Hardware Production Support Engineer... 
    Senior
    Temporary work
    Work experience placement
    Interim role
    Work at office
    Remote work
    Relocation
    Flexible hours
    Day shift

    Raytheon

    Tewksbury, MA
    23 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Staff Data Center Operations Engineer, GPU Hardware Architecture. Be the first to apply!