Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Platform Engineer: GPU Inference & Training

Perplexity

Perplexity seeks an experienced platform engineer to own and evolve a self-serve GPU compute platform. You will design and operate GPU provisioning, cluster configuration, and cross-cloud orchestration to support both training and real-time inference workloads. You’ll implement fault-tolerant scheduling, multi-cluster management, and robust observability. Candidates should have deep Kubernetes skills, GPU infra experience, and strong systems fundamentals. #J-18808-Ljbffr Perplexity

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Platform Engineer: GPU Inference & Training in San Francisco, CA vacancy
  •  ...Manufacturing Co is seeking a Software Engineer for their Supercomputing Platform & Infrastructure team. You will...  ...computing clusters, working closely with training and inference teams. The ideal candidate has experience with production GPU deployments, strong software... 
    Training

    Dormont Manufacturing Co

    San Francisco, CA
    1 day ago
  • $200k

    Platform Engineer - Inference Optimization We build and operate large-scale LLM inference and training infrastructure serving millions of users. This role focuses on deep optimization...  ...performance bottlenecks. Build and maintain GPU-native Kubernetes infrastructure,... 
    Training

    Kaon (prev. FlowGPT)

    San Francisco, CA
    1 day ago
  • Together AI in San Francisco is seeking a Research Engineer to help build a platform that lets users customize open-source models, bridging post-training and production inference. You will contribute across Fine-Tuning, RL, and Evaluation services and collaborate with product... 
    Training
    Full time
    Worldwide

    Together AI

    San Francisco, CA
    1 day ago
  •  ...infrastructure. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large‑scale training, serving, evaluation, and...  ...systems, Kubernetes, GPU infrastructure, high‑performance inference, and enterprise AI platforms to... 
    Training

    Seekr

    San Francisco, CA
    1 day ago
  •  ...out into multiple AI inference requests running in real...  ...that sits a large GPU fleet spread across several...  ...Today, our inference engineers and researchers build...  ...responsibilities we want a dedicated platform team to own. Your job...  ...platform for running training and inference... 
    Training
    Shift work

    Apply

    San Francisco, CA
    3 days ago
  •  ...company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on... 
    Training

    Reflection AI

    San Francisco, CA
    3 days ago
  • $150k - $300k

     ...that enables anyone to create, train, and deploy them. We...  ...our Solutions Architect for GPU Infrastructure, you'll be the...  ...strategies for LLM training, inference, and HPC workloads Present architectural...  ...with our world‑class engineering team while having direct impact... 
    Training

    Prime Intellect

    San Francisco, CA
    2 days ago
  •  ...s everything around it. Training runs that fail halfway through. GPU clusters that sit underutilised...  ...heavily in the platform that enables researchers...  ...ML infrastructure, and inference systems—building reliable...  ...Scientists and Research Engineers, solving the engineering... 
    Training

    techire ai

    San Francisco, CA
    1 day ago
  • $150k - $215k

    Principal Observability Platform Engineer - Nscale About Nscale Nscale is the GPU cloud engineered for AI. We provide cost‑effective, high‑performance...  ...observability: GPU utilisation, training job visibility, inference latency. Prior experience defining observability... 
    Training
    Flexible hours

    Nscale

    San Francisco, CA
    1 day ago
  • About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance...  ...The Role As a Staff Observability Platform Engineer, you'll play a critical...  ...observability challenges related to model training, inference workloads, GPU utilization, and... 
    Training

    Nscale

    San Francisco, CA
    1 day ago
  • Dormont Manufacturing Co is seeking a Staff ML Platform Engineer to join their team. This role focuses on building infrastructure for large-scale training of ML models, requiring deep knowledge of PyTorch and multi-node systems. Join a dynamic environment poised for growth... 
    Training

    Dormont Manufacturing Co

    San Francisco, CA
    1 day ago
  • $200k - $300k

     ...AI research to systems engineering to product design —...  ...looking for a Research Platform Engineer to build the...  ...Do Design and build training infrastructure, data infrastructure...  ...and own parts of the inference stack. Build internal...  ...throughput, latency, GPU utilization, and... 
    Training
    Shift work

    World Labs

    San Francisco, CA
    6 days ago
  •  ...serve humanity. We’re training and deploying frontier...  ...next generation of AI platforms powering advanced NLP...  ...for a Site Reliability Engineer to join the Model Serving...  ...with Kubernetes, and GPU workloads on those...  ...latency and throughput of inference. Strong understanding... 
    Training
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • $300k

     ...startup building an AI and cloud platform, powered by thousands of H100s...  ..., full-scale model training, or inference.  Our client operates high-performance GPU clusters powering some of the...  ..., tune, and operate inference engines such as vLLM, SGLang, and TensorRT... 
    Permanent employment
    Worldwide
    San Francisco, CA
    more than 2 months ago
  • A leading technology firm in San Francisco is seeking a skilled engineer to develop and optimize GPU infrastructure for AI model training. In this role, you'll lead technical decisions and work on scalable systems that enhance GPU optimization. Ideal candidates should have... 
    Training
    Internship

    Wafer

    San Francisco, CA
    1 day ago
  • Harrison Clarke is seeking a Senior Platform Engineer to take ownership of their core platform in San Francisco. This position involves designing multi-region Kubernetes clusters, managing GPU infrastructure, overseeing networking systems, and developing observability... 

    Harrison Clarke

    San Francisco, CA
    1 day ago
  • B Capital is seeking a Systems Engineer to join its Compute Platform team in San Francisco. This role involves maintaining a K8s-based platform and solving complex systems challenges, focusing on GPU infrastructures and multi-cloud environments. The ideal candidate has... 

    B Capital

    San Francisco, CA
    4 days ago
  •  ...vision. So we are working on training and scaling up multimodal...  ...into our core, high-throughput inference engine. Build robust and...  ...leverage thousands of expensive GPU resources. Design and implement...  ...compatible storage, and public cloud platforms (AWS). What Sets You Apart... 
    Training

    Luma AI

    San Francisco, CA
    3 days ago
  • $100k

     ...combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve...  ...About the role: As a Software Engineer on our Supercomputing Platform & Infrastructure team, you...  ...Experience working with production GPU deployments, data-intensive... 
    Training
    Local area
    Relocation
    Visa sponsorship

    Magic Inc

    San Francisco, CA
    1 day ago
  • $183k - $310k

     ...unexpectedly, or need to be improved, engineers rely on data to understand what...  ...The Role We're looking for a ML Platform Engineer with deep...  ...of the ML platform itself, from inference serving and pipeline orchestration to training infrastructure and evaluation frameworks... 
    Training
    Remote work

    Foxglove

    San Francisco, CA
    5 days ago
  • $300 per month

     ...intelligence. We’re crafting the engine that powers a world where...  ...teams to define the AI platform roadmap. Influence the long...  ...familiar with AI infrastructure (training, inference, ETL pipelines). Software...  ...performance optimizations on GPU systems and inference... 
    Training
    Temporary work
    Work at office

    Crusoe Energy Systems

    San Francisco, CA
    1 day ago
  •  ...leading fashion-tech AI company in San Francisco is seeking Software Engineers to build cutting-edge AI infrastructure. You will be responsible for designing scalable systems that support training and inference workflows, improving the reliability of AI systems and... 
    Training
    Internship

    SPREEAI

    San Francisco, CA
    1 day ago
  • $405k

     ...server fleet. You would be one of the first engineers on it. That means production firmware...  ...features for x86 and Arm (including GPU) platforms, from bring-up through production, using...  ...equivalent combination of education, training, and/or experience Required field of study... 
    Training
    Visa sponsorship

    Anthropic

    San Francisco, CA
    1 day ago
  • $160k - $250k

    Together AI is building the Inference Platform that brings the most advanced generative AI models to...  ...balancing across data centers and model engine pods. Develop auto‑scaling systems to...  ...orchestration is a strong plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and... 
    Full time
    Local area

    Together AI

    San Francisco, CA
    1 day ago
  • $250k

     ...infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is developing a fully...  ...looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud... 
    Training
    Permanent employment
    Remote work
    San Francisco, CA
    a month ago
  • $150k - $300k

    Join to apply for the Platform Engineer — AI Infra Specialist role at Poly Join to apply for the Platform Engineer — AI Infra Specialist...  ...experience Core AI principles and fundamentals around training, inference, loss, gradient descent, embeddings, tokenization and a strong... 
    Training
    Full time

    Poly

    San Francisco, CA
    2 days ago
  • $165k - $242k

     ...the most powerful end-to-end platform to develop, deploy, and iterate...  ...for how AI is built, trained, and scaled. The integration...  ...clusters, agent building, and inference at scale, we’re combining forces...  ...Role As a Senior Data Platform Engineer, you will: Partner with other... 
    Training
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Weights & Biases

    San Francisco, CA
    1 day ago
  • $295k - $405.5k

    About this role As the Senior Staff Machine Learning Platform Engineer, you will own the technical vision and evolution of Faire’s ML...  ...the long‑term architecture of Faire’s ML platform including training, inference, feature management, governance Establish company‑wide... 
    Training
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire Inc

    San Francisco, CA
    1 day ago
  • $245k - $295k

     ...We are seeking a Senior Manager, Infrastructure Platform Engineering to lead a team building core systems that turn large...  ...AKS) Familiarity with the operational challenges of GPU clusters, AI training, and inference workloads Working knowledge of platform security... 
    Training
    Temporary work
    Immediate start

    Crusoe

    San Francisco, CA
    25 days ago
  • $190k - $215k

     ...customers running production AI inference workloads and...  ...Inference, model serving platforms, LLM deployments, AI agents, or GPU‑based inference environments...  ...with the product and engineering teams to define customer...  ...infrastructure trends. Customer Training: Deliver training... 
    Training
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Platform Engineer: GPU Inference & Training. Be the first to apply!