Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Compute Cluster

Causal Labs

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it. To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather. Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems. Everything we do — training, evaluation, serving — runs on our GPU fleet. Your mission is to design, build, and operate the supercomputing environment underneath it all, delivering performant, reliable, and cost‑efficient compute to ensure research is able to iterate rapidly at scale. Responsibilities Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning Extend scheduling and orchestration systems (e.g. Kubernetes, Slurm) for topology‑aware placement, preemption, quotas, and multi‑tenancy across training and inference workloads Build software that abstracts cluster management and presents a unified, self‑serve interface to researchers and engineers Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage Monitor and continuously improve reliability and error recovery; build the observability to catch failures before researchers do Partner with researchers to unblock large‑scale runs and advise on performance and placement trade‑offs What we're looking for We value a relentless approach to problem‑solving, rapid execution, and the ability to quickly learn in unfamiliar domains. Experience operating large‑scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker) Strong systems background: Linux, networking, storage, infrastructure‑as‑code Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings Understanding of monitoring, logging, observability, and version control best practices for ML systems Familiarity with CUDA/NCCL and performance profiling for distributed workloads Owns deliverables end‑to‑end, from requirements through autonomous execution #J-18808-Ljbffr Causal Labs

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff — Compute Cluster in San Francisco, CA vacancy
  •  ...large physics foundation models for causal intelligence and weather prediction. You will design, build, and operate large-scale GPU clusters that power training, evaluation, and serving infrastructure for the research team. What You\'ll Do Design, deploy, and operate... 
    Suggested

    Linuxcareers

    San Francisco, CA
    5 days ago
  •  ...ambitious AI team. Our platform, Lab, unifies compute, environments, evaluations, secure...  ...into superintelligence they own. Core Technical Responsibilities This hybrid role spans...  ...in open development and encourage team members to contribute to the broader AI community... 
    Suggested

    Prime Intellect

    San Francisco, CA
    4 days ago
  •  ...Character.AI, Anthropic and beyond. About the Role Reflection’s Compute Platform team specializes in keeping our compute layer healthy...  ...health checks, and remediation strategies. What You’ll Do Cluster Management: Build and maintain tools for the automatic remediation... 
    Suggested
    Full time
    Relocation package

    B Capital

    San Francisco, CA
    4 days ago
  • Member of Technical Staff - Computational Biologist Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Computational Biologist Valthos Inc. Valthos is an applied biological intelligence company. We build and deploy software and biological AI systems to... 
    Suggested
    Full time
    Work at office

    Valthos

    San Francisco, CA
    1 day ago
  • $150k - $300k

     ...aggregate and orchestrate global compute into a single control plane...  ...our RL training stack. Core Technical Responsibilities LLM Serving...  ...and cold‑start times across clusters. Inference Optimization & Performance...  ...and encourage team members to contribute to the broader... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    4 days ago
  • Member of the Technical Staff, Product (Backend)Location: North America Remote / San Francisco, CA · Full...  ...reserved for hyperscalers. The first cluster filled almost instantly. The years...  ...since went into the platform that makes compute liquid: it deploys into foreign... 
    Remote work

    Andromeda

    San Francisco, CA
    3 days ago
  • $200k - $300k

     ...inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems...  ...and optimize high-performance computing kernels, and work across inference engine...  ...Design, deploy, and operate heterogeneous clusters across vendors Own customer accounts... 
    H1b
    Work at office

    Simplify

    San Francisco, CA
    2 days ago
  • $250k

     ...your career? Join a fast-growing AI compute platform building the next generation...  ...opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's...  ...stack, orchestration, scheduling, cluster management, and the agentic operations... 
    Full time
    San Francisco, CA
    10 days ago
  • Member of Technical Staff - Infrastructure Security We're partnering with a frontier AI research company...  ...containerized workloads and AI compute environments Develop automated tooling...  ...Deep expertise securing Kubernetes clusters and containerized workloads Strong experience... 

    Xcede

    San Francisco, CA
    3 days ago
  • Member of the Technical Staff, PlatformLocation: North America Remote / San Francisco, CA · Full-TimeAbout...  ...reserved for hyperscalers. The first cluster filled almost instantly. The years...  ...since went into the platform that makes compute liquid: it deploys into foreign... 
    Remote work

    Andromeda

    San Francisco, CA
    3 days ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe...  ...ultra-long context, and inference-time compute to achieve this goal. About the role...  ..., and operational clarity across clusters spanning thousands of GPUs. Magic’s long... 
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    1 day ago
  •  ...possible in robotic intelligence. As a Member of Technical Staff, you'll be at the forefront of...  ...You'll have access to Amazon's vast computational resources, enabling you to tackle ambitious...  ...scientists Leverage our massive compute cluster and extensive robotics infrastructure... 
    Local area

    Amazon Science

    San Francisco, CA
    2 days ago
  • $150k - $300k

     ...aggregate and orchestrate global compute into a single control plane...  ...-tuning runs on managed GPU clusters with a single API call or a...  ...infrastructure that runs the jobs. Core Technical Responsibilities Hosted...  ...and encourage team members to contribute to the broader... 
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Kubelt

    San Francisco, CA
    5 days ago
  • $200k - $400k

     ...possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will build...  ...going, and others still bringing up cluster nodes or deleting the third redundant...  ...idle. We run across more than one compute provider, so keeping the whole stack portable... 
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    4 days ago
  •  ...accessible to all. ABOUT THE ROLE Reflection's Compute Platform team keeps our compute layer healthy...  ...a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments... 
    Work at office
    Visa sponsorship

    B Capital

    San Francisco, CA
    22 hours ago
  • $150k - $250k

     ...Ryan Hoover (Founder, Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former Chair, Huntington Bank Ventures...  ...in investment banking or consulting BA/BS or MA/MS in Computer Science, Physics, Mathematics, or another quantitative science... 
    Full time
    Work experience placement
    Internship
    Worldwide

    Krew

    San Francisco, CA
    more than 2 months ago
  •  ...Member of Technical Staff @ Lotus AI Who we are Lotus AI is a groundbreaking primary care app that integrates your medical records, AI...  ...Experience with real-time voice AI systems, speech models, computer vision, medical imaging, or multimodal models that combine... 

    Lotus Health AI, Inc

    San Francisco, CA
    4 days ago
  • $120k - $170k

     ...with teams across the United States to help them hire. Member of Technical Staff (Founding Engineer) Location: San Francisco, CA (North...  ...Copilot, Claude Code, or similar Bachelor's degree in Computer Science or related technical field Graduated from a strong... 
    Work at office
    Remote work
    Visa sponsorship
    Relocation package

    Recruiting from Scratch

    San Francisco, CA
    3 days ago
  • $120k - $300k

     ...(Seed) · Industry: FinTech The Role A backend-leaning Member of Technical Staff building the end-to-end systems that power the client's production...  ..., self-healing pipelines); browser automation, computer vision, LLM systems. Requirements Experience with backend... 
    Full time
    H1b
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    21 days ago
  • $250k

     ...· 1–10 people (Seed) · Industry: AI Tools The Role A Member of Technical Staff who can handle everything from modeling to systems to product...  ...& APIs. Requirements A Bachelor's or Master's in Computer Science, Engineering, or a related field Strong proficiency... 
    Full time
    H1b
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    20 days ago
  • $150k - $350k

     ...develop training approaches, run experiments on distributed GPU clusters, and evaluate results. You will design and build generative...  ...structurally valid, and genuinely novel About You You have a PhD in computer science, machine learning, physics, mathematics, or a related... 

    Output Biosciences

    San Francisco, CA
    19 days ago
  • $200k - $300k

     ...with high agency. The Role Nimble is looking for a Member of Technical Staff to help us advance our robotics moonshot by designing,...  ...production systems Qualifications P.h.D in Robotics or Computer Science Experience training deep learning models for... 
    Local area
    Immediate start
    Flexible hours
    Weekend work

    Nimble Robotics

    San Francisco, CA
    3 days ago
  •  ...Perplexity is seeking an intrepid, polymathic Member of Technical Staff to take on one of the AI industry’s most unique engineering roles. You...  ...growing portfolio of frontier AI products (such as Comet and Computer). Contribute to regulatory filings and responses that... 

    aijoblist

    San Francisco, CA
    2 days ago
  • $200k - $300.09k

     ...governance structures that shaped the open web and modern computing infrastructure, the Foundation serves as a neutral, multi‑...  .... The Role The OpenClaw Foundation is seeking exceptional Members of Technical Staff (MTS) to serve as full‑time maintainers, builders, and stewards... 
    Full time

    The OpenClaw Foundation

    San Francisco, CA
    4 days ago
  •  ...patterns for both batch and streaming workloads. Build scalable compute and storage foundations (formats, engines, runtimes) that...  ...systems. Cost & Performance Management: Partitioning strategies, clustering, cost optimization, SLA-driven pipelines. About You Strong data... 
    Work at office
    Visa sponsorship

    Reflection

    San Francisco, CA
    2 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering...  ...around AI ops? How do we orchestrate heterogeneous compute resources (CPU and GPU) efficiently? What data model will enable... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    4 days ago
  •  ...About the job Member of Technical Staff – Full Stack / AI Systems Company: AdsGency AI Location: Onsite (San Francisco City) Employment Type:...  ...added by the job poster 3+ years of work experience with Computer System Design 3+ years of work experience with Python (Programming... 
    Full time
    Work experience placement
    Visa sponsorship

    Mosaic

    San Francisco, CA
    1 day ago
  • $240k - $280k

     ...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how...  ...$259,000 2 weeks ago Senior / Staff Software Engineer - Computational Chemistry / Molecular Dynamics Sr. Software Engineer -... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    3 days ago
  •  ...limiting factor in science is not data or compute. It is judgment. Our work is aimed at...  ...production deployment and iteration. Act as a technical resource within the team, mentoring...  ...training infrastructure (multi-node GPU clusters, distributed training). Experience integrating... 
    Live in

    Autopoiesis Sciences

    San Francisco, CA
    1 day ago
  • $200k

     ...ABOUT ABUNDANT AI models rely on two fundamental ingredients: compute and data. Abundant is building the NVIDIA of training data...  ...frontier startups and F500 enterprises. THE ROLE As the Member of the Technical Staff, you are the “PM of the model,” architecting the next... 
    Flexible hours

    Abundant Energy Inc.

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Compute Cluster. Be the first to apply!