Support Engineer, GPU Infrastructure
Hydra Host, Inc.
**Support Engineer, GPU Infrastructure (Tier 2/3)**Full-time | Support Operations | Reports to the Operations and Support LeadLocation: Remote**About Hydra Host**Hydra Host sells production-ready bare metal GPU compute for AI training, inference, enterprise workloads, and high-performance computing. Our control plane, Brokkr, connects customers to capacity across a network of partner-owned data centers rather than facilities we own ourselves.When a customer's workload degrades, the cause sits somewhere across their code, our platform, the facility, or the hardware vendor. Most support roles decide what is broken. Here you also decide whose it is, on live production hardware, with a customer waiting and a partner relationship on the other side of the answer.**The role**You are the person a customer issue reaches when it is real. You take it from the first symptom to a resolution that holds, working from the Linux host down through the hardware and out to the facility floor.The tier in the title is deliberate. Nearly everything that arrives at support here is already a Tier 2 or Tier 3 problem, on production hardware, with a customer's workload affected.Most of the day is diagnosis, coordinating the people with hands on the hardware, and writing down what you found so the next person does not start over. Some of it crosses into infrastructure engineering work.Support at Hydra is being built right now and you are one of the first hires into it. Expect less structure than you are used to, and more influence over what the structure becomes.**What you will do****Diagnose across the stack*** Work Linux server issues end to end. Boot and network boot failures, kernel and driver problems, filesystems, storage pressure, services, memory and CPU behavior, general instability.* Diagnose hardware failures using out-of-band management, sensor data, POST and boot errors, SMART data, and vendor diagnostics. IPMI, Redfish, iDRAC, iLO, or equivalent.* Isolate server-side network problems. NICs and drivers, VLANs, addressing, routing, MTU, DNS, DHCP, bonding, link state, and packet captures when it comes to that.* Troubleshoot NVIDIA GPU servers. GPU availability, thermal throttling, driver and VBIOS mismatch, PCIe, XID errors, and the host-level conditions that look like GPU problems and are not.* Separate hardware from OS from network from application from configuration before escalating. Being wrong about which one it is costs more than being slow.**Own the incident, not just the ticket*** Take an issue through to resolution or to a clean handoff, including the last step. A node that is repaired but still cordoned is not resolved.* Set severity by blast radius and communicate it. One node and one cluster are different events.* Notice when several tickets are one problem. Adjacent addresses, shared symptom, same time window.* Escalate to engineering with an evidence pack rather than a description, so the receiving engineer starts from your work instead of repeating it.* Take part in root cause analysis and post-incident review, and turn the findings into something that changes.**Work the partner and vendor boundary*** Drive issues with data center partners. Remote hands, reboots, cabling and optics checks, component replacement, physical inspection.* Open and track hardware RMAs with OEMs through to a replacement in the rack.* Validate repaired or replaced equipment before it returns to production.* Keep the asset record accurate. When the register and the floor disagree, close the gap rather than working around it.* Support server turn-ups, migrations and decommissions where support is involved, and validate readiness before a machine carries a customer workload.**Communicate with customers*** Write clear updates to technically sophisticated customers who want cause and timeline, not reassurance.* Ask for diagnostic information in a way that never implies the customer has misread their own situation.* Deliver an unwelcome answer plainly when that is the honest one.**Document what you learn*** Write the runbook after you solve something the first time, not the fifth.* Use accurate categories and real closure reasons. This data is how the team learns what is actually breaking.* Improve the runbooks and operational procedures you inherit as you use them.**Automation and continuous improvement*** Build scripts and small tools that take repetitive diagnostic and support work off the queue.* Use Python, Bash, or similar to automate health checks, data collection, and routine operations.* Improve the troubleshooting tools and workflows you inherit rather than working around them.* Convert recurring manual procedures into documented ones, then into automated ones.* Contribute to infrastructure-as-code and configuration management where it touches support work.* Help improve monitoring and alerting. Alerts that fire in the wrong place, or do not fire at all, are a support problem before they are anyone else's.* Work with engineering to find the changes that make the platform easier to operate and cheaper to support at scale.**Coverage**Support runs across time zones and this role works a set schedule.* A defined shift, agreed before you start.* An escalation rotation for high severity issues outside your shift hours.* A written handoff at the end of every shift. What is open, what you tried, what to pick up first.Your time zone matters to this hire. We will be direct in the first conversation about the hours we need covered.**What we are looking for*** Three or more years supporting production servers, data center infrastructure, or bare metal and cloud environments.* Strong hands-on Linux troubleshooting. You are comfortable on a console with logs, dmesg, systemd, storage tooling, and network utilities.* Real experience with server hardware. CPU and memory, storage and filesystems, RAID, PCIe, NICs, power, BIOS and UEFI, firmware and drivers.* Out-of-band management experience. IPMI, Redfish, iDRAC, iLO, or similar.* Working TCP/IP knowledge and the ability to prove whether a problem is on the host or on the network.* You reason in fault domains. How much is broken, does it survive a rebuild, does it follow the workload to another machine, can it be fixed remotely.* Experience working in ticketing, monitoring, incident management, or infrastructure management systems.* Clear written English. Most of this job happens in writing, to customers, to partners, and to engineers.* Sound judgment alone in production. You will make calls at hours when nobody is available to check them.**Helpful, not required*** NVIDIA GPU servers at scale, and any of CUDA, NCCL, NVLink, or DCGM.* HPC or AI training environments. InfiniBand or high performance Ethernet.* Enterprise platforms from Dell, HPE, Supermicro, or Lenovo.* NVMe, ZFS, Ceph, or distributed storage.* Prometheus, Grafana, or similar observability tooling.* NetBox or another infrastructure and asset register.* Ansible, Terraform, or configuration management. Git-based infrastructure workflows.* Optics, transceivers, and DAC or AOC cabling.* Working across geographically distributed third-party facilities.**What success looks like****By 90 days*** Working the majority of your shift's issues without escalation.* Every ticket you close carries an accurate category and a real closure reason.* At least three runbooks written from issues you personally resolved.* You know how to escalate to each facility contact on your shift without asking.**By six months*** Your escalations to engineering arrive complete and are rarely handed back.* Issues on your shift are increasingly caught before the customer reports them.* Something that used to be a recurring ticket is gone because you removed the cause. #J-18808-Ljbffr Hydra Host, Inc.
- Hydra Host is seeking a Support Engineer for GPU infrastructure (Tier 2/3) to diagnose issues across Linux hosts, hardware, and network layers in a production environment. This remote role involves coordinating hands-on hardware steps and writing up findings to prevent...SuggestedRemote jobWork at office
- ...Senior Infrastructure Support Engineer Chicago, Illinois, USA; Dallas, Texas, USA; Miami, Florida, USA; Philadelphia, Pennsylvania, USA; Portland, Oregon, USA As a senior Infrastructure Support Engineer, you play a vital role in maintaining technical excellence and...SuggestedWork experience placement
- ...Job Title Help Desk Support Engineer Location Doral, FL 33122 US (Primary) Category Intelligence Job Type Full-Time Career Level Staff Education Associate Degree Travel Security Clearance Required None Job Description Prescient...SuggestedFull timeRemote work
- ...Senior Level Network Support Engineer We are currently seeking a Senior Level Network Support Engineer to work in this fast-paced technology environment with world-class engineers and customers on cutting-edge technologies. Work on-site at a Customer in Miami, FL....SuggestedRemote work
- Biscayne Staffing in the United States is seeking a Senior Technical Support Specialist to own and resolve complex technical issues in a... ...troubleshoot, reproduce, document findings, collaborate with Engineering and Product, and drive customer-impacting problems to...Suggested
- Bakerly in Miami, FL is seeking an IT Administrator to support our end-user technology environment, focusing on L1/L2 helpdesk, account management, and device lifecycle tasks. This entry-to-mid-level role partners with our IT team to keep colleagues productive. You will...
- Addigy Inc. is seeking a Support Engineer, Tier 2. This remote US role focuses on resolving complex issues across Apple devices, working with Tier 1 and cross-functional teams to deliver real solutions. You will help grow the knowledge base, mentor newer engineers, and...Remote job
- ...Overview With Cyclops now actively supporting client integrations, technical questions... ...already prepared. As the Technical Support Engineer , you will own the technical support... ...is the first stablecoin and crypto infrastructure platform built exclusively for the payments...Temporary workFor contractorsRemote work
- CHA Consulting, Inc. is seeking a CEI Project Administrator/Project Engineer to join our Construction Inspection Team in Florida. You will guide inspection activities on major infrastructure projects, coordinate between owners, contractors, and staff, and ensure compliance...For contractors
- ...Role As a Software Engineer, Backend and Infrastructure, you will build the mission-critical backend powering our medical AI platform used by healthcare providers worldwide. You will architect and scale core systems, including service reliability and data platforms with...Full timeWorldwide
- Intracom Telecom is seeking a Technical Support Engineer to support world-class packet microwave and millimeter-wave radio systems used in front-haul and mobile backhaul transport for 3G/4G/5G networks, as well as fixed wireless access broadband systems. The successful...
- ...in search of a new full-time Senior AV Technician. This Senior AV Technicain owns conference room technology and hands-on end user support across the firm. This person keeps every meeting room camera-ready and every global call running without a hitch, while also...Full timeWork at officeLocal areaRemote workWorldwide
$60k - $70k
...Familiarities with Network Troubleshooting (LAN/WAN/Wi-Fi) Basic understanding of Printers and peripheral devices. Provide onsite IT support for corporate office facilities. Install, Configure and maintain IT Hardware (Computers, Printers, Network devices, Audio/Video...Work at officeRemote work- Socket.dev is seeking a Support Desk Technician to deliver fast, reliable IT assistance for incidents and service requests. You will handle inquiries via phone, email, walk-ins, and self-service, ensuring timely resolution and customer satisfaction. The role requires an...
- Bakerly in Miami, FL is seeking a Help Desk Technician to provide Level 1 IT support for employees and plant staff. This entry-level role focuses on ticket resolution, SafetyChain, and Sage-based price updates in a defined change-control process. This position emphasizes...
- Flagler Technologies MSP Team in Miami, FL seeks an IT Service Technician Tier 1 to be the first point of contact for clients experiencing technical issues, via phone, email, chat, and onsite as needed. You will troubleshoot hardware and software, manage user accounts with...
- ...satisfaction. The Opportunity We are seeking a driven, people-oriented IT Engineer to join our growing organization in Coral Gables, Miami, to provide high quality and detail-oriented IT support in a fast paced start up environment. This role will work onsite in our...Full timeWork at officeRotating shift3 days per week
- ...services and trading technology is looking for an AI Infrastructure Engineer to build and operate the platform supporting its next generation of AI-powered financial... ...secure APIs for AI-powered applications, manage GPU-based workloads across development and production...Full timeWork from home
- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure... ...that powers AI platforms, GPU-accelerated workloads, large-... ..., deterministic APIs to support governed enterprise use... ...CUDA along with LLM inference engines (TensorRT-LLM), production serving...Full timeWork experience placementLive inWork at officeLocal area
- ...IT Support Engineer Symmetry IT is a leading provider of managed services, proudly serving businesses across Florida and Texas. We... ...responsible for maintaining workstations, servers, network infrastructure, hospitality systems, and user accounts while ensuring a high...Monday to FridayWeekend work
- ...for our clients, people, shareholders, partners, and communities. Visit us at . We are seeking an experienced Senior Manager, Infrastructure Maturity & Assessment to lead rigorous baseline assessments and maturity evaluations of enterprise infrastructure estates, translating...Full timeWork experience placementLive inWork at officeLocal area
- ...Job Description Job Description Position Title: Network Support and Security Technician Reports To: System Administrator,... ...Network Security Technician to safeguard our university’s Network infrastructure and digital assets, ensuring the confidentiality, integrity,...Work at officeLocal area
- ...Job Description Job Description The Support Engineer II (L2) is a key technical escalation point within our MSP, responsible for resolving intermediate to advanced issues through phone, remote, and on-site support. This role blends deep troubleshooting expertise...Work at officeRemote workShift workNight shift
- ...United States. Join a Company that Empowers you to Build your Future The Lead Onsite IT Technician provides advanced technical support and leadership for Lennar Associates. This role involves offering advice, mentorship, and training to other technicians, ensuring...For contractorsLive inLocal areaImmediate start
- ...to join our team. This role involves maintaining and supporting the data center infrastructure, including troubleshooting, hardware installations, and... ...and execute installations with guidance from the NOC or engineering team on IT devices: Physical Servers Storage &...Remote workNight shift
- Qualifications 10+ Years of experience designing, implementing, and supporting business applications Experience in Operational Support and DevOps operations Experience in Application Development in one of the following languages: .Net C#..Net Vb, PHP, Java, etc. Experience...Night shiftWeekend work
- Job Description Perform skilled infrastructure/structured cabling work in the installation, service and maintenance, repair and alteration... ...for telecom, data, security & wireless systems practices, engineering & Federal, State & local safety standards. Strong customer/...Work experience placementLocal areaImmediate start
- ...for an experienced Technical Analyst, with a background in the support and delivery of innovative solutions to clients. The services... ...a comprehensive portfolio of consulting, applications, infrastructure and business process services.NTT DATA Services, headquartered...For contractors
$122.2k - $240.5k
Position Summary Senior Network Engineer Our Enterprise Networks practice, part of Hybrid Cloud Infrastructure within AI & Engineering, helps clients architect, modernize... ...adjacent domains. Our practitioners lead and support technical teams to assess, design, and deploy...Local area- ...FloridaOnsiteFull Time$130k - $155kAbout the Role We are seeking an experienced Senior Network Engineer to support a U.S. government client's mission-critical communications and data infrastructure in the Miami, FL area. This is a fully onsite position supporting a program that...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Support Engineer, GPU Infrastructure. Be the first to apply!
- senior application support engineer Miami, FL
- IT support engineer Miami, FL
- IT engineer Miami, FL
- technical support engineer Miami, FL
- IT developer Miami, FL
- line support engineer Miami, FL
- lab support engineer Miami, FL
- support engineer Miami, FL
- support escalation engineer Miami, FL
- customer support engineer Miami, FL




