Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Compute Cluster

Causal Labs

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it. To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather. Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems. Everything we do — training, evaluation, serving — runs on our GPU fleet. Your mission is to design, build, and operate the supercomputing environment underneath it all, delivering performant, reliable, and cost‑efficient compute to ensure research is able to iterate rapidly at scale. Responsibilities Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning Extend scheduling and orchestration systems (e.g. Kubernetes, Slurm) for topology‑aware placement, preemption, quotas, and multi‑tenancy across training and inference workloads Build software that abstracts cluster management and presents a unified, self‑serve interface to researchers and engineers Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage Monitor and continuously improve reliability and error recovery; build the observability to catch failures before researchers do Partner with researchers to unblock large‑scale runs and advise on performance and placement trade‑offs What we're looking for We value a relentless approach to problem‑solving, rapid execution, and the ability to quickly learn in unfamiliar domains. Experience operating large‑scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker) Strong systems background: Linux, networking, storage, infrastructure‑as‑code Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings Understanding of monitoring, logging, observability, and version control best practices for ML systems Familiarity with CUDA/NCCL and performance profiling for distributed workloads Owns deliverables end‑to‑end, from requirements through autonomous execution #J-18808-Ljbffr Causal Labs

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff — Compute Cluster in San Francisco, CA vacancy
  •  ...large physics foundation models for causal intelligence and weather prediction. You will design, build, and operate large-scale GPU clusters that power training, evaluation, and serving infrastructure for the research team. What You\'ll Do Design, deploy, and operate... 
    Suggested

    Linuxcareers

    San Francisco, CA
    2 days ago
  •  ...ambitious AI team. Our platform, Lab, unifies compute, environments, evaluations, secure...  ...into superintelligence they own. Core Technical Responsibilities This hybrid role spans...  ...in open development and encourage team members to contribute to the broader AI community... 
    Suggested

    Prime Intellect

    San Francisco, CA
    1 day ago
  • $200k - $400k

     ...and hardware—a position that took years to build. About the Role We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-... 
    Suggested
    Remote work
    Visa sponsorship

    Inferact Inc.

    San Francisco, CA
    3 days ago
  •  ...Character.AI, Anthropic and beyond. About the Role Reflection’s Compute Platform team specializes in keeping our compute layer healthy...  ...health checks, and remediation strategies. What You’ll Do Cluster Management: Build and maintain tools for the automatic remediation... 
    Suggested
    Full time
    Relocation package

    B Capital

    San Francisco, CA
    1 day ago
  • Member of Technical Staff - Computational Biologist Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Computational Biologist Valthos Inc. Valthos is an applied biological intelligence company. We build and deploy software and biological AI systems to... 
    Suggested
    Full time
    Work at office

    Valthos

    San Francisco, CA
    4 days ago
  • $150k - $300k

     ...aggregate and orchestrate global compute into a single control plane...  ...our RL training stack. Core Technical Responsibilities LLM Serving...  ...and cold‑start times across clusters. Inference Optimization & Performance...  ...and encourage team members to contribute to the broader... 
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    1 day ago
  • $285k - $315k

     ...building the future of high-performance compute We firmly believe that the future of...  ...possible form for whatever vendor and cluster topology you point it at, automatically...  ...production kernels. We're hiring a Member of Technical Staff for GPU Compiler Engineering to build... 
    Full time
    Work at office
    Immediate start
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    1 day ago
  •  ...Position Overview We are seeking a highly motivated & skilled Member of Technical Staff to join our growing engineering team. Member of...  ...modelling using Python & NCCL for existing & future AI compute clusters, scaling from single-GPU setups to O(100k) GPU clusters.... 
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    2 days ago
  • $275k - $315k

     ...building the future of high-performance compute We firmly believe that the future of...  ...possible form for whatever vendor and cluster topology you point it at,...  ...and safely, which is why we're hiring a Member of Technical Staff for Sandbox Infrastructure to build the... 
    Full time
    Work at office
    Relocation
    Relocation package

    SF Tensor

    San Francisco, CA
    3 days ago
  • Member of the Technical Staff, Systems Location: North America Remote / San Francisco, CA Full-Time About Andromeda Compute is the most sought-after resource in the world, yet it still trades...  ...reserved for hyperscalers. The first cluster filled almost instantly. The years... 
    Full time
    Remote work

    Andromeda

    San Francisco, CA
    1 day ago
  •  ...financial investors, distils our deep technical research and knowledge into key...  ...government agencies. Position Overview Member of Technical Staff will play a crucial role in developing...  ...& NCCL for existing & future AI compute clusters, scaling from single-GPU setups to O... 
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    3 days ago
  •  ...coordinating with external APIs, managing GPU clusters, and handling failures that...  ...Mission district office. Your Role: As a Member of Technical Staff, you will be responsible for building...  ...infrastructure/cloud resources. Computer Vision & Multimodal Media: Experience... 
    Work at office
    Immediate start
    Flexible hours
    Night shift

    Mixpeek

    San Francisco, CA
    1 day ago
  • $200k - $300k

     ...inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems...  ...and optimize high-performance computing kernels, and work across inference engine...  ...Design, deploy, and operate heterogeneous clusters across vendors Own customer accounts... 
    H1b
    Work at office

    Simplify

    San Francisco, CA
    4 days ago
  • Member of Technical Staff - Distributed Systems San Francisco, CA Full-time | On-site Join an early...  ...layer for heterogeneous AI compute across CPUs, GPUs, and emerging accelerators...  ...infrastructure beyond simply operating clusters Collaborate across inference,... 
    Full time

    Acceler8 Talent

    San Francisco, CA
    5 days ago
  • $250k

     ...your career? Join a fast-growing AI compute platform building the next generation...  ...opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's...  ...stack, orchestration, scheduling, cluster management, and the agentic operations... 
    Full time
    San Francisco, CA
    a month ago
  • $200k - $350k

    Member of Technical Staff — LLM Research & Training About the Role We are looking for an exceptional...  ...distributed training across large GPU clusters. Design and optimize model‑parallel...  ...compiler optimization, or high‑performance computing is a plus. Research Background We... 
    H1b
    Visa sponsorship

    Pragmatike

    San Francisco, CA
    2 days ago
  • Member of Technical Staff - Infrastructure Security We're partnering with a frontier AI research company...  ...containerized workloads and AI compute environments Develop automated tooling...  ...Deep expertise securing Kubernetes clusters and containerized workloads Strong experience... 

    Xcede

    San Francisco, CA
    5 days ago
  • $150k - $300k

     ...aggregate and orchestrate global compute into a single control plane...  ...-tuning runs on managed GPU clusters with a single API call or a...  ...infrastructure that runs the jobs. Core Technical Responsibilities Hosted...  ...and encourage team members to contribute to the broader... 
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Kubelt

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will build...  ...going, and others still bringing up cluster nodes or deleting the third redundant...  ...idle. We run across more than one compute provider, so keeping the whole stack portable... 
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    1 day ago
  • $275k - $315k

     ...building the future of high-performance compute We firmly believe that the future of...  ...possible form for whatever vendor and cluster topology you point it at, automatically...  ...-deployable models. We're hiring a Member of Technical Staff to own the modeling side of that end to... 
    Full time
    Work at office
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    5 days ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe...  ...ultra-long context, and inference-time compute to achieve this goal. About the role...  ..., and operational clarity across clusters spanning thousands of GPUs. Magic’s long... 
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    4 days ago
  •  ...accessible to all. ABOUT THE ROLE Reflection's Compute Platform team keeps our compute layer healthy...  ...a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments... 
    Work at office
    Visa sponsorship

    B Capital

    San Francisco, CA
    3 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering...  ...around AI ops? How do we orchestrate heterogeneous compute resources (CPU and GPU) efficiently? What data model will enable... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    1 day ago
  •  ...About the job Member of Technical Staff – Full Stack / AI Systems Company: AdsGency AI Location: Onsite (San Francisco City) Employment Type:...  ...added by the job poster 3+ years of work experience with Computer System Design 3+ years of work experience with Python (Programming... 
    Full time
    Work experience placement
    Visa sponsorship

    Mosaic

    San Francisco, CA
    3 days ago
  • $240k - $280k

     ...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how...  ...$259,000 2 weeks ago Senior / Staff Software Engineer - Computational Chemistry / Molecular Dynamics Sr. Software Engineer -... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    5 days ago
  • $2,000 per month

     ...faster than anyone in the world. Job Summary We’re hiring Members of Technical Staff — the general research/engineering seat. If you have a seriously...  ..., and more Daily lunch and dinner in our office Unlimited compute budget subject to ROI justification Unlimited Codex and... 
    Work at office
    Relocation package

    Build AI

    San Francisco, CA
    4 days ago
  •  ...environments. What you'll do as a Member of Technical Staff - Infrastructure at Phylo: Design...  ...services, sandboxed agent execution, compute workloads, storage mounts, networking...  ...processing, or scientific compute clusters. Experience with enterprise... 
    Work at office

    Phylo

    South San Francisco, CA
    23 hours ago
  • $180k - $250k

     ...frontier AI can actually do computational biology. Pharma, biotech, and...  ...Postgres and blobs on S3 Cluster orchestrator designed to accept...  ...where engineers talk through technical challenges and stay aligned on...  ...• Founders & Chief of Staff: Alfredo, Kyle, Kenny, Jordan... 
    Work at office
    Visa sponsorship
    Work visa

    LatchBio

    San Francisco, CA
    5 days ago
  • $200k - $300k

     ...with high agency. The Role Nimble is looking for a Member of Technical Staff to help us advance our robotics moonshot by designing,...  ...production systems Qualifications P.h.D in Robotics or Computer Science Experience training deep learning models for... 
    Local area
    Immediate start
    Flexible hours
    Weekend work

    Nimble Robotics

    San Francisco, CA
    5 days ago
  •  ...Member of Technical Staff @ Lotus AI Who we are Lotus AI is a groundbreaking primary care app that integrates your medical records, AI...  ...Experience with real-time voice AI systems, speech models, computer vision, medical imaging, or multimodal models that combine... 

    Lotus Health AI, Inc

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Compute Cluster. Be the first to apply!