Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff GPU Compute Infrastructure Engineer

causal

Causal is building a Large Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate large distributed GPU clusters. You will extend schedulers, build self-serve interfaces, and own storage and lineage for checkpoints and logs. You will collaborate with researchers to optimize performance and placement, ensuring reliable, scalable compute for rapid research iteration. #J-18808-Ljbffr causal

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff GPU Compute Infrastructure Engineer in San Francisco, CA vacancy
  • B Capital is seeking a Systems Engineer to join its Compute Platform team in San Francisco. This role involves maintaining a K8s-based platform...  ...and solving complex systems challenges, focusing on GPU infrastructures and multi-cloud environments. The ideal candidate has... 
    Suggested

    B Capital

    San Francisco, CA
    2 days ago
  •  ...applied AI research, flexible infrastructure, and seamless developer...  ...and help build the platform engineers turn to to ship AI products....  ...workloads scale, the network is the computer. We are looking for foundational engineers to lead our GPU Networking efforts, making RDMA... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    13 hours ago
  • $188k - $275k

     ...CoreWeave combines superior infrastructure performance with deep technical...  ...breakthroughs and turn compute into capability. Founded in...  ...What You'll Do: The Field Engineering organization at CoreWeave is...  ...customer lifecycle: leading new GPU cluster bring-up and... 
    Suggested
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    28 days ago
  •  ...these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands...  ...matters to the world. The Production Engineering Team Examples of key exciting...  ...10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput... 
    Suggested
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    13 hours ago
  • $230k - $405k

    About the Team:Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize...  ...scale networking protocols, RDMA, NCCL, GPU hardware behavior, benchmarking,... 
    Suggested
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • StratITech is hiring a senior network engineer to architect and operate secure, high-performance data center and edge networks supporting AI and distributed compute workloads. You will translate loosely defined needs into concrete requirements and plans, and own the end... 

    StratITech

    San Francisco, CA
    4 days ago
  • Nscale seeks a Senior Infrastructure Support Engineer to own the health of GPU fleets and high‑performance fabrics. You will operate across GPU hardware, Linux, and data centre operations, bridging Support, DC Operations, and Engineering. You’ll diagnose complex issues... 
    Remote work

    Nscale

    San Francisco, CA
    1 day ago
  • $350k

    Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized,... 
    Visa sponsorship

    Thinking Machines Lab

    San Francisco, CA
    1 day ago
  • Thinking Machines Lab Inc. is seeking a network engineer to own the lowest layers of the network stack for large-scale training and inference. You will ensure interconnect reliability across GPU fabrics, debugging NICs, and building instrumentation for faster troubleshooting... 

    Thinking Machines Lab Inc.

    San Francisco, CA
    4 days ago
  •  ...We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations...  ...between our researchers and the bare GPU machines, helping to make sure that...  .... You're not afraid of physical computers We’re building out edge datacenters... 
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    25 days ago
  • $148.7k - $201.2k

     ...Silicon (MQS) Center for Quantum Computing (CQC) is a multi-disciplinary team of scientists, engineers, and technicians, on a...  ...-performance computing (HPC) infrastructure on AWS that CQC scientists and...  ...from the needs of our science staff in the context of our larger... 
    Local area
    Flexible hours

    Amazon

    San Francisco, CA
    2 days ago
  • $167k - $230k

     ...psychological safety, empathy, and human connection—that will allow employees of all backgrounds to thrive.Senior Software Engineer - Analytics Compute Platform TeamAbout the Role & TeamEvery chart, insight, and experiment result a customer sees in Amplitude passes... 
    Work at office
    Home office
    Flexible hours

    Amplitude

    San Francisco, CA
    4 days ago
  •  ...medical products.Are you a seasoned platform engineer passionate about cloud-native technologies...  ...? Do you thrive on solving complex compute challenges by leveraging the power of Kubernetes? Join our dynamic Infrastructure Platform team and help us drive a critical... 
    Local area
    Worldwide
    Relocation

    HeartFlow

    San Francisco, CA
    2 days ago
  • Crusoe is seeking a Staff TPM to own deployment programs for new sites and capacity expansions in our GPU‑powered AI cloud. You will define targets, gating criteria, and DRI matrices, coordinating chip vendors and OEMs while managing parallel capacity projects. You will... 

    Crusoe

    San Francisco, CA
    13 hours ago
  •  ...Senior Infrastructure Engineer Vast.ai's cloud powers AI projects and businesses all over the world...  ...democratizing and decentralizing AI computing—reshaping our future for the benefit of...  ...systems that power Vast.ai's global GPU marketplace. You'll work closely with... 
    Full time
    Work at office

    Vast

    San Francisco, CA
    3 days ago
  •  ...quickly eclipsing physics-based tools in computational drug discovery. Scientists often...  ...the Role We're looking for two Infrastructure Engineers to lead the scaling of our machine learning...  ...tools Experience with GPU workloads Technology Our technology... 
    Relocation

    Tamarind Bio, Inc

    San Francisco, CA
    4 days ago
  • $188.28k - $270.38k

     ...Staff Computational Biologist Remote At Freenome, we are seeking a seasoned Staff Computational Biologist to help grow the Freenome Computational Science team. The ideal candidate is an expert in computational biology, genomics, and modeling, with a proven track... 
    Local area
    Remote work

    Freenome

    Brisbane, CA
    4 days ago
  • Prime Intellect, Inc. seeks a Compute Strategy Lead to oversee GPU sourcing, economics, and contracts. This role involves shaping the industry by negotiating substantial agreements and collaborating with research teams on compute requirements. Key qualifications include... 
    Remote job
    Flexible hours

    Prime Intellect, Inc.

    San Francisco, CA
    4 days ago
  •  ...ABOUT THE ROLE Reflection's Compute Platform team keeps our compute...  ...a team of strong systems engineers, guide the technical and architectural...  ..., and next-generation GPU deployments, and work closely...  ...mentoring, and growing systems or infrastructure teams while staying... 
    Work at office
    Visa sponsorship

    B Capital

    San Francisco, CA
    1 day ago
  •  ...6505687Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA...  ...a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks...  ...large-scale workloads across extensive GPU clusters.Build developer... 
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    13 hours ago
  •  ...· 5+ years of hands-on experience in infrastructure engineering with demonstrated depth and breadth of...  ...review, and documentation. · Degree in Computer Science, Computer Engineering,...  ...(Intel Xeon, AMD EPYC, NVMe storage, GPU accelerators) sufficient to evaluate hardware... 
    Temporary work

    PB consulting

    San Francisco, CA
    26 days ago
  • Hyperbolic is seeking a very senior Infrastructure Engineer to scale a GPU cloud marketplace. You will build a multi-tenant provisioning and virtualization layer, turning global GPU inventories into an orchestrated pool for AI developers and researchers. You will own the... 

    Hyperbolic

    San Francisco, CA
    3 days ago
  • A tech company specializing in AI infrastructure is seeking a Software Engineer to build a scalable compute platform for its generative video models. The ideal candidate will have over 5 years of experience in MLOps or AI infrastructure management, along with strong Python... 

    HeyGen

    San Francisco, CA
    4 days ago
  • $350k

     ...the Role We're looking for a network engineer to own the lowest layers of the network...  ...interconnect reliability at scale, across large GPU fabrics - both the RDMA/RoCE fabric...  ...Bachelor's degree or equivalent experience in computer science, engineering, or similar.... 
    Immediate start
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    4 days ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...Lambda's mission is to make compute as ubiquitous as electricity...  .... One person, one GPU. If you'd like to build the...  ...on-call rotation for Network Engineering team You Have 10+... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    2 days ago
  •  ...Network Engineer Gimlet Labs is seeking a Network Engineer to design...  ...build, and scale the network infrastructure powering production-scale AI...  ..., and other high-performance compute environments. This is an...  ...Experience with AI/HPC, GPU, or large-scale distributed infrastructure... 

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • $190k - $280k

     ...Senior Network Engineer San Francisco About the Role Together...  ...operate the global network infrastructure supporting our production...  ...services and high-performance AI compute environments. This is a...  ...fabrics. Experience supporting GPU clusters, HPC environments,... 
    Full time

    Together AI

    San Francisco, CA
    4 days ago
  • $150k - $250k

     ...demanding. The Role As a Network Engineer, Autonomous Factory, you’ll...  ...the full network stack: physical infrastructure, switching and routing, server and GPU rack fabric, segmentation between...  ...systems, fieldbus segments, edge compute, cameras, and operator stations.... 
    Full time
    For contractors
    Local area
    Remote work

    Foundry Robotics

    San Francisco, CA
    3 days ago
  • Luma AI in SF Bay Area is hiring for a Research Scientist/Engineer in Training Infrastructure. You will design distributed training systems for thousands of GPUs and enable researchers to push multimodal foundation models. You should have deep experience with PyTorch, CUDA... 
    Remote job

    Luma AI

    San Francisco, CA
    1 day ago
  • $7.5k

     ...Job Description Job Description Infrastructure Engineer Location: San Francisco, CA (In-Office) Partnership: EQL Tech has been exclusively...  ...as a Quant at Goldman Sachs, and the other built the computer vision system for the largest smart warehousing company globally... 
    Work at office
    Relocation
    Visa sponsorship
    Relocation package
    Day shift

    EQL Tech

    San Francisco, CA
    10 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff GPU Compute Infrastructure Engineer. Be the first to apply!