Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Engineer, Distributed GPU Clusters

Kindredventures

Kindredventures is building a Large Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute environment. You will enable research at scale by provisioning, upgrading, and optimizing distributed clusters that power training and inference workloads. You will extend orchestration, implement topology-aware scheduling, and deliver a self-serve platform for researchers, with strong focus on reliability, observability, and cost efficiency. #J-18808-Ljbffr Kindredventures

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff Engineer, Distributed GPU Clusters in San Francisco, CA vacancy
  • Causal Labs is building a Large Physics foundation Model and GPU-driven compute environment to enable rapid research iteration at scale. You will design, deploy, and operate massive GPU clusters, extending Kubernetes and Slurm for efficient, multi-tenant workloads. You... 
    Suggested

    Causal Labs

    San Francisco, CA
    1 day ago
  • Causal is building a Large Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate large distributed GPU clusters. You will extend schedulers, build self-serve interfaces, and own storage and lineage for checkpoints and logs. You will... 
    Suggested

    causal

    San Francisco, CA
    4 days ago
  • Magic AI, Inc. is seeking a engineer for the Supercomputing Platform & Infrastructure to design, build, and operate large-scale GPU infrastructure powering model training and inference...  ...hybrid environments, manage Kubernetes clusters, and ensure reproducibility and... 
    Suggested
    Visa sponsorship
    Relocation package

    Magic AI Corp.

    San Francisco, CA
    2 days ago
  •  ...firm in San Francisco is seeking a candidate to build and scale distributed training systems for large model pre-training. You will...  ...experience with modern distributed training frameworks and optimizing GPU utilization. This role offers competitive compensation and a supportive... 
    Suggested

    Reflection

    San Francisco, CA
    1 day ago
  • $192k - $260k

     ...leading data and AI company is seeking a Staff Engineer to design and implement core systems...  ...of experience in building large-scale distributed systems and will collaborate closely across...  ...to ensure operational excellence in GPU serving workloads. Competitive salary range... 
    Suggested

    Databricks Inc.

    San Francisco, CA
    4 days ago
  • $168k - $247k

     ....About the RoleAs a Senior/Staff Deep RL Engineer, you will design, train, and...  ...design through large-scale distributed training to on-vehicle...  ...based deep RL agents using GPU-accelerated simulation at massive...  ...JAX across large compute clusters.Build agentic optimization... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    5 hours ago
  • $173.5k - $331.05k

     ...video rendering pipeline behind the next generation of these products. We are looking for a senior, hands-on engineer to own and evolve the cross-platform GPU rendering platform at the heart of that mission — leading major initiatives through to completion in close partnership... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    3 days ago
  • $179k - $218k

     ...Silicon Reality" must be bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical...  ..., and predictive strategies needed to maintain peak cluster health.The Strategic BridgeFor DC Engineering: You are... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $250k - $300k

     ...Staff Engineer, Distributed Storage and HPC & AI InfrastructureSan FranciscoAbout the RoleIn this role,...  ..., low-latency data paths, and massive cluster-scale storage operations.You will also...  ...architectural decisions as we scale our GPU fleet.Engineer and scale multi-... 
    Full time
    Remote work

    Together AI

    San Francisco, CA
    3 days ago
  •  ...company in San Francisco is seeking a Member of Technical Staff focused on kernels and GPU performance. This role involves optimizing GPU and...  ...various hardware. Ideal candidates have strong software engineering foundations and experience with performance-critical systems... 

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • B Capital is seeking a skilled engineer for GPU infrastructure in San Francisco. This role involves designing and operating high-performance systems for model inference, synthetic data generation, and reinforcement learning. The ideal candidate has strong GPU systems experience... 

    B Capital

    San Francisco, CA
    5 days ago
  • $150k - $350k

    Gimlet Labs, Inc. is seeking a Member of Technical Staff focused on optimizing GPU and accelerator kernels for AI workloads. This role involves...  ...Ideal candidates possess a strong foundation in software engineering and experience with performance-critical systems. Compensation... 

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  •  ...Token Company in San Francisco is seeking a Member of Technical Staff for their infrastructure team. In this role, you will own the cloud...  ...compression API and build global low-latency, high-throughput GPU ML inference infrastructure. The ideal candidate will have solid... 
    Visa sponsorship

    The Token Company

    San Francisco, CA
    3 days ago
  • $209k - $253k

     ...our Compute-focused Production Engineers are the backbone of that...  ...and HPC workloads across CPU, GPU, and DPU/NIC resources. You will...  ...or maintaining custom Linux distributions or kernels for specific platforms...  ...orchestration across GPU clusters.Contributions to Linux kernel... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $300 per month

     ...This Role:We are seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s...  ...reliability across Crusoe Cloud’s GPU- and CPU-based...  ...and platform insights into cluster-level tuning and configuration...  ....Hands-on experience with distributed training and/or inference... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • Wafer is building AI-powered GPU optimization systems and is seeking engineers to join a small, highly collaborative team. You will work directly with the founders to implement the agent framework, profiling, and compiler tooling that power our GPU optimization platform... 

    Wafer

    San Francisco, CA
    1 day ago
  •  ...financial world.The roleWe are seeking a Staff Vulnerability Management Engineer to lead the most complex technical...  ...hardware-adjacent surfaces such as GPU, DPU/BlueField, BMC, and firmware....  ..., including cloud, containers, and distributed systems.Strong programming or... 
    Remote work

    SoFi

    San Francisco, CA
    1 day ago
  • $210k - $255k

     ...Crusoe.About the Role:We are looking for a Staff Engineer to be the detection authority for...  ...CCIA team and operate as a peer to the Distributed Systems Staff Engineer on the team.What...  ...systems including straggler node detection, GPU health signals, and fleet-level... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...company in San Francisco is seeking a Member of Technical Staff to design and build distributed systems for AI workloads. The role involves developing...  ...APIs. Ideal candidates should have strong software engineering skills and experience with distributed systems. This role... 

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • Eventual is seeking a Member of Technical Staff to build Eventual's core products and...  ...autonomously solving problems. We value engineers who can scope tasks and implement efficient...  ..., with a focus on performance and reliability across distributed #J-18808-Ljbffr Mixpeek

    Mixpeek

    San Francisco, CA
    4 days ago
  • $150k - $350k

    Gimlet Labs, Inc. is seeking a Member of Technical Staff to focus on distributed systems in San Francisco, California. This role involves designing...  ...APIs. The ideal candidate should have strong software engineering fundamentals and experience with distributed systems. The... 

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  • Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient...  ...our systems. The ideal candidate has a strong background in distributed systems and is eager to engage in complex challenges. Enjoy a... 

    Sail Research

    San Francisco, CA
    4 days ago
  • $156k - $190k

     ...us at Crusoe.About the Role:As a Staff Cloud Support Engineer, you are a technical authority...  ...orchestration (Slurm, Terraform), and AI/ML cluster stability.Reduce MTTR and...  ...ExpertiseTroubleshoot NCCL, IB, GPU driver/firmware issues, distributed training failures.Support... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $250k - $300k

     ...Crusoe.About the Role:As a Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be...  ...of large-scale, multi-node GPU clusters. You will own the CI/CD infrastructure...  ...from low-level Linux Systems up to Distributed Control Planes.Working knowledge... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $207k - $290k

     ...are seeking an experienced AI Engineer with deep expertise in...  ...to join our team as a Senior Staff Architect. In this role, you...  ...Scalability & Infrastructure: Design distributed training systems, leverage...  ...(Kubernetes, GPU/TPU clusters, and cloud ML platforms).... 
    Worldwide
    Flexible hours

    JazzX AI

    San Francisco, CA
    more than 2 months ago
  • $190.9k - $232.8k

    A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of...  ...engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong software... 

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $220k - $280k

     ...Staff MLOps Engineer — Machine Learning Platform Location: New York, NY / San Francisco, CA / Remote...  ...of machine learning, backend distributed systems, and platform performance. You...  ...trade-offs between execution latency, GPU/CPU throughput, and cloud infrastructure... 
    Remote work

    GrabJobs

    San Francisco, CA
    4 days ago
  •  ...stake real consequences on. As a Staff Machine Learning Engineer, you’ll own AI-driven products end to...  ...something reliable at scale, drawing on real distributed-systems experience. You’re energized...  ...concurrency inference (Triton, vLLM, GPU-backed serving) that stays fast and... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Primer.ai

    San Francisco, CA
    10 hours ago
  •  ...are seeking an experienced AI Engineer with deep expertise in...  ...to join our team as a Senior Staff Architect. In this role, you...  ...Scalability & Infrastructure: Design distributed training systems, leverage...  ...infrastructure (Kubernetes, GPU/TPU clusters, and cloud ML platforms).... 
    Flexible hours

    JazzX AI

    San Francisco, CA
    2 days ago
  • Harrison Clarke is working with a high-growth startup in San Francisco seeking a Staff Distributed Systems Engineer. This hands-on role involves designing and building core systems for a cutting-edge AI code generation product. The ideal candidate will have excellent coding... 

    Harrison Clarke

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Engineer, Distributed GPU Clusters. Be the first to apply!