Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Engineer - Large-Scale GPU Inference & RL Infra

Visa Hunt

Reflection is a research lab focused on making intelligence open and accessible. We design, build, and operate GPU-heavy infrastructure for high-throughput model inference and mid-training workloads. Join a team that powers synthetic data generation, RL pipelines, and distributed model evaluation across thousands of GPUs. Expect collaboration with research teams and a focus on performance tuning, kernel optimization, and scalable distributed systems. #J-18808-Ljbffr Visa Hunt

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Engineer - Large-Scale GPU Inference & RL Infra in San Francisco, CA vacancy
  • B Capital is seeking a skilled engineer for GPU infrastructure in San Francisco. This role involves designing and operating high-performance systems for model inference, synthetic data generation, and reinforcement learning. The ideal candidate has strong GPU systems experience... 
    Suggested

    B Capital

    San Francisco, CA
    1 day ago
  • Causal Labs is building a Large Physics foundation Model and GPU-driven compute environment to enable rapid research iteration at scale. You will design, deploy, and operate massive GPU clusters, extending Kubernetes and Slurm for efficient, multi-tenant workloads. You... 
    Suggested

    Causal Labs

    San Francisco, CA
    2 days ago
  • $192k - $260k

     ...leading data and AI company is seeking a Staff Engineer to design and implement core systems...  ...over 10 years of experience in building large-scale distributed systems and will collaborate...  ...to ensure operational excellence in GPU serving workloads. Competitive salary range... 
    Suggested

    Databricks Inc.

    San Francisco, CA
    17 hours ago
  • Kindredventures is building a Large Physics foundation Model...  ...seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute...  ...will enable research at scale by provisioning, upgrading...  ...that power training and inference workloads. You will extend... 
    Suggested

    Kindredventures

    San Francisco, CA
    17 hours ago
  • causal is building a Large Physics foundation Model to enable general causal intelligence...  ...and identify actions to alter it while scaling novel architectures across multimodal physical...  ...data. We are seeking infrastructure engineers who can design, implement, and optimize distributed... 
    Suggested

    causal

    San Francisco, CA
    17 hours ago
  • $173.5k - $331.05k

     ...are looking for a senior, hands-on engineer to own and evolve the cross-platform GPU rendering platform at the heart of...  ...proficiency in modern C++ in a large, complex, cross-platform codebase...  ...GPU driver / hardware vendors ML inference integration (e.g., TensorRT, ONNX... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 days ago
  • $168k - $247k

     ...improve continuously from large-scale fleet data. Our models jointly...  ....About the RoleAs a Senior/Staff Deep RL Engineer, you will design, train,...  ...distributed training to on-vehicle inference. You'll help define how...  ...based deep RL agents using GPU-accelerated simulation at... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    1 day ago
  • Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work...  ...a track record in building efficient inference stacks, GPU-aware optimization, and deep learning frameworks like PyTorch... 

    Kindredventures

    San Francisco, CA
    17 hours ago
  • Fluidstack is seeking a senior network deployment engineer to lead end-to-end fabric turn-ups across data centers. You will own technical...  ...cannot close. Candidates should have proven experience on large-scale networks, automation in Python or Go, and familiarity with optics... 

    Fluidstack

    San Francisco, CA
    1 day ago
  • $220k - $280k

     ...Staff MLOps Engineer — Machine Learning Platform Location...  ...latency and enterprise scale. The world's...  ...deployment, distributed inference pipelines, and...  ...latency, GPU/CPU throughput, and...  ...serving and optimizing large language models (LLMs...  ...validate robust infra prototypes quickly... 
    Remote work

    GrabJobs

    San Francisco, CA
    4 days ago
  •  ...firm in San Francisco is seeking a candidate to build and scale distributed training systems for large model pre-training. You will collaborate with research...  ...modern distributed training frameworks and optimizing GPU utilization. This role offers competitive compensation... 

    Reflection

    San Francisco, CA
    3 days ago
  •  ...Barnes Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration,...  ...production environment. You’ll contribute to inference platforms, model serving, and high-... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    3 days ago
  • $273k - $345k

     ...until they work at scale. We are roboticists, engineers, operators, and builders...  ...cutting edge RL and distillation techniques...  ...multimodal systems, large-scale VLA models, or...  ...Profile real-time inference pipelines to...  ...and eliminate CPU, GPU, and memory bandwidth... 
    Full time
    Internship
    Work at office
    Flexible hours

    Atoms

    San Francisco, CA
    2 days ago
  • United States Digital Space LLC is seeking an infrastructure leader to own a self-serve GPU compute platform for training and inference workloads. You will design and operate the system that lets researchers launch jobs across multi-cloud GPU fleets without manual provisioning... 

    United States Digital Space LLC

    San Francisco, CA
    17 hours ago
  • $210k - $255k

     ...urgency, who believe in the scale of our ambition and...  ...:We are looking for a Staff Engineer to be the detection authority...  ...node detection, GPU health signals, and fleet...  ...as the fleet grows.ML/RL Integration: Evaluate and...  ....Familiarity with large-scale fleet management... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $250k - $300k

     ...urgency, who believe in the scale of our ambition and thrive on...  ...will spend your time making large language models run faster, cheaper...  .... That means owning the inference stack end to end: profiling where...  ...work directly with customer engineering teams to tailor deployments... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $179k - $218k

     ...sense of urgency, who believe in the scale of our ambition and thrive on a...  ...be bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive...  ...DCGM and ROCm). Experience using large datasets or basic ML frameworks to... 
    Temporary work

    Crusoe

    San Francisco, CA
    17 hours ago
  • $207k - $290k

     ...enterprises don't scale expertise—they...  ...experienced AI Engineer with deep expertise...  ...Learning (RL) to join our team as a Senior Staff Architect. In this...  ...designing and deploying large-scale RL...  ...techniques , including inference-time search,...  ...(Kubernetes, GPU/TPU clusters, and... 
    Worldwide
    Flexible hours

    JazzX AI

    San Francisco, CA
    more than 2 months ago
  •  ...visual-text dataset to train large-scale models from scratch that are continuously...  ...pod is a small group (~10 engineers) inside Labs, which allows for...  ...function development for RL training.   What you’ll do...  ...experience with large-scale GPU computing. ~ Build flexible... 
    Full time
    Work experience placement
    Currently hiring
    Work at office
    Remote work
    Relocation
    Relocation package
    Flexible hours

    Pinterest

    San Francisco, CA
    17 hours ago
  •  ...consequences on. As a Staff Machine Learning Engineer, you’ll own AI-driven...  ...across the modern stack: large language models, agentic...  ...into something reliable at scale, drawing on real...  ...latency, high-concurrency inference (Triton, vLLM, GPU-backed serving) that stays... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Primer.ai

    San Francisco, CA
    17 hours ago
  • $197.3k - $313.7k

     ...Slack is looking for a Staff Machine Learning Engineer with deep expertise...  ...finetuning strategies for large language models and...  ...pipelines on GPU infrastructure.Brainstorm...  ...serve real users at scale — not just research prototypes...  ...optimization for inference (quantization,... 
    Full time

    Salesforce

    San Francisco, CA
    2 days ago
  •  ...Job Description Staff Machine Learning Engineer, Artificial Intelligence...  ...evaluation systems, inference architecture, and...  ...deployment, including GPU optimization, memory...  ...latency reduction, and scaling policies. - Collaborate...  ...working with large models and understanding... 
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    28 days ago
  • $300 per month

     ..., who believe in the scale of our ambition and thrive...  ...:We are seeking a Staff Hardware Systems Engineer to strengthen Crusoe’...  ...bring-up to large-scale production while...  ...across Crusoe Cloud’s GPU- and CPU-based infrastructure...  ...across training and inference - dense, MoE, long-... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...The role: SoFi’s Staff AI Engineer is a hands-on AI engineering...  ...reasoning problems at scale. Distributed Agent...  ..., low-latency inference across diverse hardware...  ...Deep understanding of Large Language Model (LLM) architectures...  ...underlying Kubernetes/GPU orchestration for... 
    Full time

    Sofi

    San Francisco, CA
    17 hours ago
  • $200k - $350k

     ...We are seeking a Staff AI Engineer to lead the design and...  ...production AI systems at scale. What You'll Do...  ...Design model serving, inference, evaluation, and optimization...  .... Kubernetes and GPU infrastructure. RAG...  ...on complex systems, large-scale integrations, and... 
    Remote job
    Full time
    Immediate start

    Pragmatike

    San Francisco, CA
    8 hours ago
  •  ...embodied AI. We are hiring a Member of Technical Staff to lead reinforcement learning infrastructure...  ...work with Qi and the research team to improve large vision-language and code-generating agents through fine-tuning, RL, and scalable evaluation in a on-site San Francisco... 

    Doist

    San Francisco, CA
    17 hours ago
  • $227.33k - $312.58k

    We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier...  ...that power frontier‑scale machine learning...  ...transform, and operate on large‑scale datasets that...  ...training, evaluation, or inference workflows.Solid...  ...Optimizing data pipelines for GPU‑backed training and... 
    Full time
    Work at office
    Local area
    Immediate start
    3 days per week

    Procore Technologies

    San Francisco, CA
    4 days ago
  • $190.9k - $232.8k

    A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong... 

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $156k - $190k

     ..., who believe in the scale of our ambition and thrive...  ....About the Role:As a Staff Cloud Support Engineer, you are a technical...  ...by preventing large-scale incidents. You...  ...ExpertiseTroubleshoot NCCL, IB, GPU driver/firmware...  ...(training + inference) with performance tuning... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...financial world.The role: We’re seeking a Staff Security Detection Engineer to build and mature SoFi’s machine...  ..., and validation, operating over large-scale security data lakes and streaming...  ...training, feature pipelines, and inference.MLOps practices - feature stores, model... 
    Remote work

    SoFi

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Engineer - Large-Scale GPU Inference & RL Infra. Be the first to apply!