Get new jobs by email
  • Inferact Inc. is building a world-class GPU compute platform powering vLLM. We seek a hands-on cluster administration engineer to own the high-performance infrastructure that keeps our engineers productive. You will monitor GPU servers, manage scheduling with SLURM/Kubernetes... 
    Suggested
    Remote work

    Inferact Inc.

    Brooklyn, NY
    4 days ago
  • $44.5k - $51.5k

     ...Medical, Dental, Vision ~ Life Insurance ~ Long-term/Short-term Disability ~ Accident Insurance ~ Critical Insurance Our Cluster Sales Managers make a difference by: A positive outlook and outgoing personality A team-first attitude A gift for paying... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Courtyard Noblesville

    Noblesville, IN
    1 day ago
  •  ...for a Frontend Engineer to own and shape the interface of our cluster operating system — translating complex distributed state into...  ...Kubernetes offerings. Experience creating visualization systems and compute-efficient UI animations. Familiarity with analytics tools... 
    Suggested
    Full time

    Clera

    San Francisco, CA
    1 day ago
  •  ...iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises...  ...software systems that deploy, validate, and manage AI compute clusters across data centers worldwide for the world’s fastest AI... 
    Suggested
    Full time
    Worldwide

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • SpaceX is hiring a Sr. HPC Systems Engineer in Hawthorne, CA to administer HPC clusters, storage systems, and fast networks. You will provide application support to engineers, install Linux compute clusters, and document procedures for diverse teams, conveying complex... 
    Suggested

    SPACE EXPLORATION TECHNOLOGIES CORP

    Hawthorne, CA
    4 days ago
  •  ...Systems Engineer for the Massachusetts Green High Performance Computing Center (MGHPCC). You will support the computing infrastructure...  ...responsibilities include deploying, maintaining, and optimizing HPC clusters and storage systems dedicated to AI/ML workloads. This hands-... 
    Suggested

    Massachusetts Institute of Technology

    Cambridge, MA
    4 days ago
  •  ...challenges at scale, and helping to create safer, more civil shared experiences for everyone. As a Principal Software Engineer on the Compute Cell Lifecycle team you will create, support, and evolve the infrastructure at Roblox as we build out Roblox's private cloud. The... 
    Suggested
    Full time

    Roblox

    Remote
    1 day ago
  •  ...during the training of the frontier models. About the Role We are looking for engineers to operate the next generation of compute clusters that power OpenAI’s frontier research. This role blends distributed systems engineering with hands-on infrastructure work... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...Job Description You will need to manage windows servers and be experience on windows servers, i.e. Windows 2k12 (VM's) setup a cluster, domain controllers, web servers, updates and security. IT or System Administrator needed for related issues System or IT... 
    Suggested
    Full time

    Mapjects.com

    Remote
    1 day ago
  • $207k - $275k

     ...deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded...  ...2025. Learn more at  . About the role As part of the Cluster Orchestration team, you will play a key role in advancing... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Core Weave

    Remote
    1 day ago
  •  ...iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises...  ...inference.The RoleWe are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge... 
    Suggested

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • POSITION SUMMARYInstall, configure, manage, maintain, test, evaluate, and repair computer networks, workstations, support server system(s), supporting hardware/software, user accounts, and computer/telephone rooms. Train/instruct users in proper use and security of all... 
    Suggested
    Work experience placement

    Marriott International

    Melbourne, FL
    1 day ago
  •  ...looking for exceptional people to join us! About the role The Compute team's mission: any engineer, using AI, should be able to stand...  ...change safety matter Nice to haves Deep K8s internals: cluster lifecycle, admission controllers, custom controllers/operators... 
    Suggested
    Full time
    For contractors
    Internship

    Persona

    San Francisco, CA
    1 day ago
  • $241k - $331k

     ...understand why disease happens and how to correct it. With our compute capacity, AI research and engineering, and state-of-the-art technology...  ...our understanding of human health. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform... 
    Suggested
    Full time
    Work at office
    Worldwide
    Relocation package
    3 days per week

    Biohub

    Remote
    1 day ago
  •  ...high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server...  ...networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.Experience designing, scaling, or operating... 
    Suggested

    AMD

    San Jose, CA
    7 hours ago
  • $153.2k - $254.5k

     ...significant visibility within the organization and aligns with our mission of becoming the world's most customer-centric company. The DCEO Cluster Manager position encompasses leadership responsibility for multiple data center facilities and their corresponding infrastructure.... 
    For subcontractor
    Work at office
    Flexible hours
    Shift work

    Amazon

    Fairless Hills, PA
    1 day ago
  •  ...solutions, Render offers a developer-first experience with persistent compute, dynamic autoscaling, built-in orchestration, and observability...  ...Render builds and orchestrates a growing number of kubernetes clusters on different hyperscalers and recently, our own hardware. We... 
    Full time
    Remote work
    Worldwide
    Home office

    Render

    United States
    1 day ago
  • $180k

     ...accurately share knowledge with their teammates. About the Role The Compute Infrastructure team at xAI is responsible for designing, building, and operating the massive-scale clusters and orchestration platforms that power frontier AI training, inference, and agent... 
    Full time
    Temporary work

    Xai

    Palo Alto, CA
    1 day ago
  • $145.92k - $209.24k

     ...merchant supplier - delivering integrated quantum solutions across computing, networking, sensing, and security. IonQ’s newest generation of...  ...international. Job ID: 1750The Role: We're looking for an HPC Cluster Engineer to join our Infrastructure Team. Our mission is to... 
    Permanent employment
    Contract work
    Work at office
    Remote work

    IONQ

    College Park, MD
    1 day ago
  • $175k - $300k

     ...building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human...  ...Migrate live compute at construction speed: we're converting clusters across production sites simultaneously, bringing new sites online... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    1 day ago
  • $320k - $405k

     ...imperative to optimize how we use it. As a Software Engineer for Compute Efficiency on the Capacity team, you will play a central role...  ...cloud service providers and internal stakeholders to optimize cluster configurations, workload placement, and resource utilization across... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Remote
    1 day ago
  • $176k - $276k

     ...looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building...  ...an outstanding engineer, be a key player to the most exciting computing hardware and software to contribute to the latest breakthroughs... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $122k - $167.75k

     ...cyber security technologies More about this roleWhat you need to know about this position:What extra ingredients you will bring:The Cluster Automation & Controls Lead is responsible for manufacturing automation systems and cybersecurity in terms of maintaining a constant... 
    Full time
    Local area
    Relocation package
    Monday to Friday

    Mondelēz International

    Richmond, VA
    2 days ago
  •  ...powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors,...  ...serve as the technical program leader for large-scale compute cluster bring-up, driving execution from planning through delivery. You... 
    Work at office

    AMD

    Texas
    2 days ago
  • $160k - $275k

     ...performant and functionally accurate silicon for MatX products across compute, memory management. High-speed connectivity and other key...  ...systems usable: from Linux kernel drivers up through node and cluster management. The team also co-owns the BMC/OpenBMC firmware stack... 
    Daily paid
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    2 days ago
  • $175k - $287k

     ...team. LinkedIn operates one of the largest privately managed compute infrastructures in the world outside the public cloud providers...  ...problems that span large-scale workload orchestration, multi-cluster operations, platform reliability and efficiency, GPU compute optimization... 
    Full time
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Remote
    1 day ago
  •  ...machines. And we're only getting started. At Databricks, the Compute Infrastructure organization builds and operates the foundation...  ...tens of millions of VMs per day, operates thousands of Kubernetes clusters, and must deliver extreme elasticity, reliability and cost... 
    Full time

    Databricks

    Remote
    1 day ago
  •  ...with Vera Rubin on the horizon — on billions of dollars worth of compute, in collaboration with partners that are the largest public AI...  ...and you want to spend the next few years making the most expensive GPU clusters on the planet earn their keep, we'd love to talk.... 
    Full time
    Work at office
    Flexible hours
    Night shift

    Eventual

    Remote
    1 day ago
  • $132.3k - $198.45k

     ...looking for senior engineers to build/scale Nuro's large-scale computing infrastructure in the cloud/data center. This system is the...  ...orchestrate and execute large-scale workloads in cloud and on-premise clusters. Collaborate with application teams throughout Nuro to... 
    Full time

    Nuro

    Mountain View, CA
    1 day ago
  • $165k - $225k

     ...performance AI infrastructure for organizations running intensive computational research, large-scale model training, and demanding data...  ...and implement the compute orchestration layer that manages GPU clusters, bare-metal provisioning, and resource scheduling-enabling... 
    Full time
    Immediate start
    Remote work
    Flexible hours

    Moon Lite, Inc.

    Chicago, IL
    1 day ago