Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Systems Engineer — Distributed Training & Inference

TensorScale AI

TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical improvements that scale. Responsibilities include CUDA / Triton optimizations, designing efficient distributed inference and training pipelines, and owning communication performance across GPUs and nodes with #J-18808-Ljbffr TensorScale AI

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the GPU Systems Engineer — Distributed Training & Inference in San Francisco, CA vacancy
  •  ...in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional... 
    Training

    Baseten

    San Francisco, CA
    1 day ago
  •  ...excellence. We seek engineers with strong intrinsic...  ...We’re looking for a systems engineer with HPC or...  ...experience to help scale AI inference. You’ll leverage your...  ...systems to optimize GPU performance at the...  ...Familiarity with distributed training/inference frameworks... 
    Training
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    4 days ago
  •  ...of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution...  ...will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-... 
    Training

    Magic AI, Inc

    San Francisco, CA
    1 day ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 
    Suggested

    Vast.ai Inc.

    San Francisco, CA
    4 days ago
  •  ...San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will...  ...ensure efficient hardware utilization across multi‑node GPU clusters. Join a team focused on scalable AI foundations... 
    Training

    Genesis AI

    San Francisco, CA
    2 days ago
  • $200.8k - $251k

     ...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    14 hours ago
  •  ...Baseten powers mission-critical inference for the world's most...  ...and help build the platform engineers turn to to ship AI products...  ...building the global operating system for distributed, heterogeneous AI hardware....  ...engineers to lead our GPU Networking efforts, making... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  • $250k

     ...building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company...  ...Site Reliability Engineer to support and scale...  ...observability across distributed compute environments...  ...available infrastructure systems Improve CI/CD... 
    Training
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...Francisco is seeking a Staff ML Systems Engineer to design and prototype...  ...low-latency, high-throughput inference. You will implement changes...  ...systems, while profiling across GPU, networking, and memory to...  ...also co-design RL and post-training pipelines, drive performance... 
    Training

    Together

    San Francisco, CA
    14 hours ago
  •  ...AI Systems Engineer - Codex Core Agents Location San Francisco...  ..., model behavior, inference/runtime issues, and product...  ...systems in distributed systems, infrastructure...  ...model evals, or post‑training feedback loops. Background...  ...optimization, GPU systems, benchmarking... 
    Training
    Full time
    Work at office
    Local area
    Relocation package
    Flexible hours

    Slope

    San Francisco, CA
    4 days ago
  •  ...BeamBeam is an ultrafast AI inference platform. We built a...  ...runtime that launches GPU-backed containers in...  ...help us with Platform Engineering work. We're working on...  ...problems:Low-level systems development: working with...  ...working with a large distributed systemComfortable with... 

    BEAM inc.

    San Francisco, CA
    5 days ago
  • $300 per month

     ...seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s Hardware...  ...across Crusoe Cloud’s GPU- and CPU-based infrastructure...  ...studies across training and inference - dense, MoE, long-context...  ....Hands-on experience with distributed training and/or inference... 
    Training
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $350k

     ...and steerable AI systems. We want AI to be...  ...committed researchers, engineers, policy experts,...  ...Anthropic's inference fleet serves Claude...  ...kernels, model servers, distributed routing,...  ...systems, especially training or inference infrastructure...  ...Familiarity with GPU/TPU/accelerator... 
    Training
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    2 days ago
  •  ...foundation Model and seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute environment. You will...  ...provisioning, upgrading, and optimizing distributed clusters that power training and inference workloads. You will extend orchestration,... 
    Training

    Kindredventures

    San Francisco, CA
    1 day ago
  • $225k

     ...Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and...  ...high reliability for RL and post-training workflows. The ideal candidate will possess... 
    Training

    Dormont Manufacturing Co

    San Francisco, CA
    14 hours ago
  •  ...design, build, and operate GPU-heavy infrastructure for high-throughput model inference and mid-training workloads. Join a team that...  ...generation, RL pipelines, and distributed model evaluation across thousands...  ..., and scalable distributed systems. #J-18808-Ljbffr Visa Hunt
    Training

    Visa Hunt

    San Francisco, CA
    2 days ago
  • $150k - $300k

     ...anyone to create, train, and deploy them....  ...Solutions Architect for GPU Infrastructure,...  ...production‑ready systems capable of...  ...for LLM training, inference, and HPC workloads...  ...and Kubernetes for distributed workloads Implement...  ...with our world‑class engineering team while having... 
    Training

    Prime Intellect

    San Francisco, CA
    14 hours ago
  •  ...in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible...  ...ideal candidate has a strong background in distributed systems and is eager to engage in... 

    Sail Research

    San Francisco, CA
    1 day ago
  •  ...is recruiting infrastructure engineers to scale large-scale inference and evaluation around a...  ...will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes...  ...efficient inference stacks, GPU-aware optimization, and deep... 

    Kindredventures

    San Francisco, CA
    1 day ago
  • $227.2k - $417k

     ...the Role:As a Software Engineer on the ML...  ...class machine learning inference platforms. These platforms...  ...latency ML model serving systems that support Deep Learning...  ...throughput, and low latency distributed systems using...  ...), ElastiCache, model training orchestration, etc.Understanding... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    4 days ago
  •  ...’re hiring a hands‑on Vision Systems Engineer to own the detection, tracking...  ...Discrimination: Develop and train lightweight CNN classifiers for...  ...for real‑time embedded inference via quantization, pruning, or...  ...detect — on embedded FPGA and GPU platforms Performance Modeling... 
    Training

    Exploration Technology Group

    San Francisco, CA
    5 days ago
  • $150k - $300k

     ...enables anyone to create, train, and deploy them. We...  ...LLM serving, LLM inference optimization and RL systems. You will be working...  ...across our cloud GPU fleets. GPU‑Aware Scheduling...  ...SLOs. Model Distribution: Optimize model...  ...PyTorch: LLM Inference engine development and integration... 
    Training
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    1 day ago
  •  ...an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting...  ...will work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding... 
    Training

    Mixpeek

    San Francisco, CA
    3 days ago
  •  ...technologies—like the da Vinci surgical system and Ion—have transformed how...  ...worldwide.We’re a team of engineers, clinicians, and innovators...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you...  ...models to real-time onboard inference—while serving as a core... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    10 hours ago
  • $160k - $194k

     ...organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of...  ...level, skill sets, market conditions, experience and training, licensure and certifications, and business and... 
    Training
    Full time
    Remote work
    Flexible hours

    Ad Verse

    San Francisco, CA
    14 hours ago
  •  ...Fluidstack Production Engineering TeamExamples of key exciting problems...  ...actual state inspection, and distributed command execution. One...  ...not a hundred scripts.Make the system's view of itself always match...  ...platforms, so every new site and GPU generation lands cleanly from... 
    Local area

    Fluidstack

    San Francisco, CA
    5 days ago
  • $110 per hour

     .... Position: MLOps Engineer (JAX, PyTorch, Pallas/...  ...performance in MLOps , training infrastructure, and...  ...to MLOps and ML systems problems . Evaluate...  ...training pipeline design, distributed systems reasoning, and...  ...or optimizing custom GPU kernels using Pallas... 
    Training
    Remote job
    Contract work
    Summer work
    Weekday work

    Mercor

    San Francisco, CA
    5 days ago
  •  ...company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on... 
    Training

    Reflection AI

    San Francisco, CA
    1 day ago
  • $117.2k - $223.9k

     ...experiences. Join our team of talented engineers and help us advance the...  ...infrastructure security, and the security of distributed and scalable distributed systems. This role requires hands-on...  ...compensation, promotion, benefits, training, assessment of job performance, discipline... 
    Training
    Full time

    Salesforce

    San Francisco, CA
    4 days ago
  •  ...practical constraints of robotic platforms. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines,... 
    Training
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    14 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Systems Engineer — Distributed Training & Inference. Be the first to apply!