Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

HPC & GPU Infrastructure Support Engineer

$90k - $160k

Vast.ai

About Us Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation. We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state‑of‑the‑art AI systems while collaborating with a globally distributed team. About the Role This role focuses on troubleshooting complex Linux and GPU infrastructure issues across NVIDIA drivers, CUDA, GPU workloads, Ubuntu, Docker, KVM‑based virtual machines, networking, hardware, BIOS, and firmware. You’ll investigate failures, reproduce issues, identify root causes, and propose practical solutions across the full infrastructure stack. You’ll also serve as the engineering resource our L1 support team relies on when tickets go beyond frontline triage. You’ll own complex escalations end‑to‑end, gather technical evidence, coordinate with the appropriate teams, and communicate findings clearly to clients, infrastructure suppliers, and internal teams. The best engineers in this role don’t just resolve individual issues—they recognize recurring patterns, improve diagnostic tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic Linux, GPU, and infrastructure problems. Strong GPU troubleshooting experience, Linux systems knowledge, and technical support skills are the primary requirements. You should be comfortable working autonomously in Ubuntu environments and troubleshooting NVIDIA drivers, CUDA, containers, virtual machines, networking, hardware, and GPU workloads. Vast.ai users or hosts strongly preferred. Location and Schedule This is a full‑time position based in our Westwood, Los Angeles office. Available schedules: Monday–Friday: Fully on‑site Sunday–Thursday: Four days on‑site and one day working from home Key Responsibilities Diagnose and resolve issues across NVIDIA CUDA/GPU drivers, Docker, and KVM virtualization environments Investigate GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU‑accelerated workloads Troubleshoot network‑layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host‑side configuration Provide managed support for supplier onboarding and ongoing machine management, including installation, configuration, and post‑setup troubleshooting Advise suppliers on hardware setup, driver configuration, BIOS and firmware settings, and network configuration for optimal performance Provide coverage for L1 support overflow during peak periods or incidents Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues You Are Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line Someone who enjoys debugging difficult problems and fixing broken systems Methodical and focused on finding root causes, not just temporary fixes Able to manage complex tickets independently A clear written communicator with an interest in AI infrastructure and GPU computing Must-Haves Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting Hands‑on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting Python and Bash scripting skills for automation and diagnostic tooling Strong written English communication that is clear, professional, and technically precise Experience providing technical support in a customer‑facing or internal help desk environment Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escape versus own resolution end‑to‑end Nice‑to‑Haves Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU‑accelerated containers Monitoring and observability experience (Prometheus, Grafana) Relevant certifications: RHCSA, CompTIA Linux+, or similar Knowledge of the Vast.ai platform as a client or infrastructure supplier Interview Process (~1 week) After you submit your application, our technical team will review your experience and qualifications. Selected candidates will proceed through the following stages: 15 minutes — Initial Screening (Virtual): A brief conversation about your background, availability, and interest in the role 45 minutes — Experience Interview (Virtual): An introduction to Vast.ai and a deeper discussion of your technical and support experience 2 hours — Meet and Greet and Technical Assessment (On‑site): Meet the team and complete an LLM‑assisted Linux systems operations assessment Annual Salary Range $90,000 – $160,000 + equity + benefits Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential. Benefits Comprehensive health, dental, vision, and life insurance 401(k) with company match Meaningful early‑stage equity Onsite meals, snacks, and close collaboration with founders/tech leaders Ambitious, fast‑paced startup culture where initiative is rewarded #J-18808-Ljbffr Vast.ai

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the HPC & GPU Infrastructure Support Engineer in Los Angeles, CA vacancy
  • $40 - $47 per hour

    iSoftStone, Inc. is seeking an Office IT Infrastructure and Support Engineer to Join our Team in Los Angeles, CA! THIS IS A PART-TIME, ONE-YEAR W2 CONTRACT POSITION WHICH REQUIRES TRAVEL BETWEEN SITES IN CARSON, IRVINE, AND LOS ANGELES CANDIDATES MUST BE ONSITE 5 DAYS... 
    Suggested
    Contract work
    Part time
    Work at office
    Worldwide

    iSoftStone

    Los Angeles, CA
    9 hours ago
  • A global IT service company in Los Angeles is seeking an Office IT Infrastructure and Support Engineer for a part-time, one-year contract. The successful candidate will be responsible for maintaining office IT infrastructure, including network troubleshooting, IDF/server... 
    Suggested
    Contract work
    Part time
    Work at office

    iSoftStone

    Los Angeles, CA
    1 day ago
  • $130k - $160k

     ...with the ultimate goal of enabling human life on Mars.IT NETWORK ENGINEER - TOP SECRET CLEARANCE SpaceX is looking for an experienced...  ...build, deployment, and operation of classified IT networks which support critical national security missions. This position will own the... 
    Suggested
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $110k - $130k

     ...Mendix, part of Siemens Digital Industries Software, is hiring an Application Support Engineer in Los Angeles, CA. This role is crucial for managing customer issues and ensuring the optimal use of the low-code platform. Candidates should possess strong technical and problem... 
    Suggested

    Mendix

    Los Angeles, CA
    2 days ago
  • Vast.ai in Los Angeles is seeking a senior technical support engineer specializing in escalated infrastructure issues. You will diagnose and resolve issues across Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM), while building runbooks and tooling to reduce repeats... 
    Suggested
    Full time
    Work at office

    Vast.ai Inc.

    Los Angeles, CA
    1 day ago
  •  ...Los Angeles is seeking an experienced Network Engineer to design, implement, and maintain the network infrastructure. This role encompasses managing everything from...  ...design and operations, particularly in systems supporting high-performance computing. Strong proficiency... 

    Generalmatter

    Los Angeles, CA
    2 days ago
  • $130k - $195k

    A leading technology firm is seeking a Technical Support Engineer in Los Angeles to enhance customer support for technical users including AI engineers. The role involves diagnosing customer issues, working closely with product teams, and mentoring support engineers. Required... 
    Remote job

    LangChain

    Los Angeles, CA
    3 days ago
  • $130k - $160k

     ...goal of enabling human life on Mars.IT INFRASTRUCTURE ENGINEER, VIRTUALIZATION & STORAGEAs an IT...  ...software lifecycle while scaling and supporting data center infrastructure, virtualization...  ...and multi-tier storage (enterprise + HPC) and be ready to build the next generation... 
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $160k - $225k

     ...NETWORK AUTOMATION ENGINEERSpaceX is looking for an experienced Sr. Network Automation Engineer to help build the tools necessary to support our large and growing network infrastructure. This employee will be a member of the IT Network Engineering team, which is... 
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    19 hours ago
  •  ...world-renowned science and engineering institute that marshals some...  ...seeking a Network Engineer to support the design, optimization,...  ...the high-performance network infrastructure powering world-class NASA, NSF...  ...high-performance computing (HPC) and research archive... 
    Permanent employment
    Casual work
    Remote work
    1 day per week

    California Institute of Technology

    Pasadena, CA
    1 day ago
  • SpaceX is seeking an IT RCDD / ICT Design Engineer to design and support critical infrastructure for production facilities, launch infrastructure and spacecraft systems. You will design fiber and copper networks, coordinate downtime, manage contracts, and support field... 

    InvestedintheMission

    Hawthorne, CA
    4 days ago
  •  ...our cutting-edge AI research. As a Data Infrastructure Engineer, you will lead the development of...  ...data infrastructure and systems needed to support our AI applications. Examples includeLarge...  ...Qualifications:Experience with GPU computingExperience with distributed data... 

    HeyGen

    Los Angeles, CA
    19 hours ago
  •  ...are seeking a seasoned Software Engineer to build and scale the foundational compute infrastructure that powers our state-of-the-art...  ...every AI-generated video.Optimize GPU Utilization: Design and...  ...scale MLOps, AI infrastructure, or HPC systems.Experience with data frameworks... 
    Full time

    HeyGen

    Los Angeles, CA
    19 hours ago
  • $90k - $150k

     ...About the Role This is a technical support role focused on escalated infrastructure issues that go beyond frontline triage. You'll be the engineering resource our L1 support team leans on...  ...networking, Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM). You'll... 
    Full time
    Work at office

    Vast.ai Inc.

    Los Angeles, CA
    3 days ago
  • Riot Games is seeking a Senior Software Engineer to enhance the Riot Client experience for players. You will deliver robust technical solutions, working closely with a diverse team in a collaborative environment. Applicants should have over 6 years of experience, a strong... 
    Flexible hours

    Riot Games

    Los Angeles, CA
    9 hours ago
  •  ...Mainframe Systems Programmer/DBA is responsible for leading and/or supporting the most complex database upgrade projects, modification,...  ...the possession of a bachelor's degree in an IT-related or engineering field. Additional Education Required Additional... 

    Trinus

    Downey, CA
    1 day ago
  • $100k - $110k

     ...Los Angeles, California, is dedicated to supporting creative professionals and advancing...  ...services.Role SummaryThis on-site Systems Engineer role in Los Angeles, California is...  ...for supporting and optimizing enterprise infrastructure across hybrid environments. Working closely... 
    Immediate start

    Robert Half

    Los Angeles, CA
    1 day ago
  • $250k - $320k

     ...define the technical direction for the infrastructure platform powering automated manufacturing...  ..., setting the standards, patterns, and engineering culture that the entire organization...  ...insurance plans for employees401kRelocation support may be provided for certain situations,... 
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    2 days ago
  •  ...programs. Our platform gives hardware engineering teams a single place to ingest data, analyze...  ....We're looking for a Baremetal Infrastructure Engineer to join our team building high...  ...infrastructure modernization, all while supporting deployments in highly constrained... 
    Permanent employment
    Work at office

    Nominal

    Los Angeles, CA
    19 hours ago
  • $140k - $160k

     ...Los Angeles and New York.About the RoleWe’re looking for an Infrastructure Engineer who loves building reliable, scalable systems. You will...  ...work closely with Engineering, Data, Product, and Design to support a fast-moving organization, and as we build products both on... 
    Work at office

    Ghost

    Los Angeles, CA
    1 day ago
  • $200k - $280k

     ...define the technical direction for the infrastructure platform powering automated manufacturing...  ..., setting the standards, patterns, and engineering culture that the entire organization...  ...insurance plans for employees401kRelocation support may be provided for certain situations,... 
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    2 days ago
  •  ...Whatnot updates on our news and engineering blogs and join us as we...  ...ll design and scale the core infrastructure that powers machine learning...  ...training & high-throughput GPU inference. What you'll do:...  ...critical business surfaces-supporting growth, recommendations, trust... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Work from home
    Home office
    Flexible hours

    Whatnot

    Los Angeles, CA
    4 days ago
  • $85k - $95k

     ...Job Description The IT Client Support Engineer serves as a frontline resource for end users at the Santa Monica location, providing dependable...  ...-leading advanced technology, data analytics and digital infrastructure and the highly rated and first-of-its-kind STARZ app.... 
    Full time

    Starz

    Santa Monica, CA
    1 day ago
  • $130k - $195k

     ...About the Role We’re hiring a Technical Support Engineer to lead our customer support experience for highly technical users, from AI engineers to infrastructure architects. You’ll be on the front lines helping teams debug production LLM applications and agents, improve... 
    Remote work

    LangChain

    Los Angeles, CA
    5 days ago
  • $160k - $270k

     ...looking for a Director, Platform Engineering to build and lead a new team...  ...AI-assisted solutions that support broadcast engineering,...  ...engineering, DevOps, MLOps, or infrastructure engineering teamsStrong...  ...developmentExperience with GPU compute, local model training... 
    Full time
    Local area

    Fox

    Los Angeles, CA
    3 days ago
  •  ...keep the world cheering.The Role AXS is seeking a Sr. Solutions Engineer, Marketing Cloud to join our Marketing Technology team. This...  ...role responsible for designing, developing, implementing, and supporting marketing automation solutions, customer journeys, and data integrations... 
    Full time
    Local area
    Worldwide
    Flexible hours

    AXS Group

    Los Angeles, CA
    1 day ago
  • $70 - $85 per hour

    DescriptionWe are seeking a Senior L4 Network Engineer / Architect with strong expertise in...  ...be designed, deployed, integrated, and supported across MPLS, Internet, broadband, LTE/5...  ..., firewalls, and LAN aggregation infrastructure. The role requires hands-on technical depth... 
    Contract work
    Temporary work
    Remote work

    TEKsystems

    Los Angeles, CA
    1 day ago
  • $165k - $230k

     ...enabling human life on Mars.SR. NETWORK ENGINEER (STARSHIELD)Starshield leverages SpaceX...  ...technology and launch capability to support national security efforts. While Starlink...  ...points of presence (POPs)Manage network infrastructure and deployments to Top Secret... 
    Permanent employment
    Temporary work
    Immediate start
    Remote work
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  • $120 - $130 per hour

     ...is seeking a highly experienced Network Engineer to perform a comprehensive assessment,...  ...optimization of its enterprise network infrastructure.This role is intended for a level Cisco...  ...Enterprise Networking• Design, implement, support, and optimize enterprise network... 
    Contract work
    Temporary work
    Interim role
    Immediate start
    Remote work

    TEKsystems

    Los Angeles, CA
    1 day ago
  • $128.5k - $298.1k

     ...Description The Cloud Engineer will design, build, and operate infrastructure and applications supporting UCLA Health's Analytics Platform across both on-premises and multi...  ...scalable training/inference platforms (GPU/accelerated workloads): capacity planning/quotas... 

    University of California

    Los Angeles, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to HPC & GPU Infrastructure Support Engineer. Be the first to apply!