Platform Engineer: GPU Inference & Training
Perplexity
Perplexity seeks an experienced platform engineer to own and evolve a self-serve GPU compute platform. You will design and operate GPU provisioning, cluster configuration, and cross-cloud orchestration to support both training and real-time inference workloads. You’ll implement fault-tolerant scheduling, multi-cluster management, and robust observability. Candidates should have deep Kubernetes skills, GPU infra experience, and strong systems fundamentals. #J-18808-Ljbffr Perplexity
- ...Manufacturing Co is seeking a Software Engineer for their Supercomputing Platform & Infrastructure team. You will... ...computing clusters, working closely with training and inference teams. The ideal candidate has experience with production GPU deployments, strong software...Training
$200k
Platform Engineer - Inference Optimization We build and operate large-scale LLM inference and training infrastructure serving millions of users. This role focuses on deep optimization... ...performance bottlenecks. Build and maintain GPU-native Kubernetes infrastructure,...Training- Together AI in San Francisco is seeking a Research Engineer to help build a platform that lets users customize open-source models, bridging post-training and production inference. You will contribute across Fine-Tuning, RL, and Evaluation services and collaborate with product...TrainingFull timeWorldwide
- ...infrastructure. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large‑scale training, serving, evaluation, and... ...systems, Kubernetes, GPU infrastructure, high‑performance inference, and enterprise AI platforms to...Training
- ...out into multiple AI inference requests running in real... ...that sits a large GPU fleet spread across several... ...Today, our inference engineers and researchers build... ...responsibilities we want a dedicated platform team to own. Your job... ...platform for running training and inference...TrainingShift work
- ...company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on...Training
$150k - $300k
...that enables anyone to create, train, and deploy them. We... ...our Solutions Architect for GPU Infrastructure, you'll be the... ...strategies for LLM training, inference, and HPC workloads Present architectural... ...with our world‑class engineering team while having direct impact...Training- ...s everything around it. Training runs that fail halfway through. GPU clusters that sit underutilised... ...heavily in the platform that enables researchers... ...ML infrastructure, and inference systems—building reliable... ...Scientists and Research Engineers, solving the engineering...Training
$150k - $215k
Principal Observability Platform Engineer - Nscale About Nscale Nscale is the GPU cloud engineered for AI. We provide cost‑effective, high‑performance... ...observability: GPU utilisation, training job visibility, inference latency. Prior experience defining observability...TrainingFlexible hours- About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance... ...The Role As a Staff Observability Platform Engineer, you'll play a critical... ...observability challenges related to model training, inference workloads, GPU utilization, and...Training
- Dormont Manufacturing Co is seeking a Staff ML Platform Engineer to join their team. This role focuses on building infrastructure for large-scale training of ML models, requiring deep knowledge of PyTorch and multi-node systems. Join a dynamic environment poised for growth...Training
$200k - $300k
...AI research to systems engineering to product design —... ...looking for a Research Platform Engineer to build the... ...Do Design and build training infrastructure, data infrastructure... ...and own parts of the inference stack. Build internal... ...throughput, latency, GPU utilization, and...TrainingShift work- ...serve humanity. We’re training and deploying frontier... ...next generation of AI platforms powering advanced NLP... ...for a Site Reliability Engineer to join the Model Serving... ...with Kubernetes, and GPU workloads on those... ...latency and throughput of inference. Strong understanding...TrainingFull timeWork experience placementWork at officeRemote workFlexible hours
$300k
...startup building an AI and cloud platform, powered by thousands of H100s... ..., full-scale model training, or inference. Our client operates high-performance GPU clusters powering some of the... ..., tune, and operate inference engines such as vLLM, SGLang, and TensorRT...Permanent employmentWorldwide- A leading technology firm in San Francisco is seeking a skilled engineer to develop and optimize GPU infrastructure for AI model training. In this role, you'll lead technical decisions and work on scalable systems that enhance GPU optimization. Ideal candidates should have...TrainingInternship
- Harrison Clarke is seeking a Senior Platform Engineer to take ownership of their core platform in San Francisco. This position involves designing multi-region Kubernetes clusters, managing GPU infrastructure, overseeing networking systems, and developing observability...
- B Capital is seeking a Systems Engineer to join its Compute Platform team in San Francisco. This role involves maintaining a K8s-based platform and solving complex systems challenges, focusing on GPU infrastructures and multi-cloud environments. The ideal candidate has...
- ...vision. So we are working on training and scaling up multimodal... ...into our core, high-throughput inference engine. Build robust and... ...leverage thousands of expensive GPU resources. Design and implement... ...compatible storage, and public cloud platforms (AWS). What Sets You Apart...Training
$100k
...combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve... ...About the role: As a Software Engineer on our Supercomputing Platform & Infrastructure team, you... ...Experience working with production GPU deployments, data-intensive...TrainingLocal areaRelocationVisa sponsorship$183k - $310k
...unexpectedly, or need to be improved, engineers rely on data to understand what... ...The Role We're looking for a ML Platform Engineer with deep... ...of the ML platform itself, from inference serving and pipeline orchestration to training infrastructure and evaluation frameworks...TrainingRemote work$300 per month
...intelligence. We’re crafting the engine that powers a world where... ...teams to define the AI platform roadmap. Influence the long... ...familiar with AI infrastructure (training, inference, ETL pipelines). Software... ...performance optimizations on GPU systems and inference...TrainingTemporary workWork at office- ...leading fashion-tech AI company in San Francisco is seeking Software Engineers to build cutting-edge AI infrastructure. You will be responsible for designing scalable systems that support training and inference workflows, improving the reliability of AI systems and...TrainingInternship
$405k
...server fleet. You would be one of the first engineers on it. That means production firmware... ...features for x86 and Arm (including GPU) platforms, from bring-up through production, using... ...equivalent combination of education, training, and/or experience Required field of study...TrainingVisa sponsorship$160k - $250k
Together AI is building the Inference Platform that brings the most advanced generative AI models to... ...balancing across data centers and model engine pods. Develop auto‑scaling systems to... ...orchestration is a strong plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and...Full timeLocal area$250k
...infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is developing a fully... ...looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud...TrainingPermanent employmentRemote work$150k - $300k
Join to apply for the Platform Engineer — AI Infra Specialist role at Poly Join to apply for the Platform Engineer — AI Infra Specialist... ...experience Core AI principles and fundamentals around training, inference, loss, gradient descent, embeddings, tokenization and a strong...TrainingFull time$165k - $242k
...the most powerful end-to-end platform to develop, deploy, and iterate... ...for how AI is built, trained, and scaled. The integration... ...clusters, agent building, and inference at scale, we’re combining forces... ...Role As a Senior Data Platform Engineer, you will: Partner with other...TrainingPermanent employmentTemporary workCasual workWork at officeFlexible hours$295k - $405.5k
About this role As the Senior Staff Machine Learning Platform Engineer, you will own the technical vision and evolution of Faire’s ML... ...the long‑term architecture of Faire’s ML platform including training, inference, feature management, governance Establish company‑wide...TrainingWork experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours3 days per week$245k - $295k
...We are seeking a Senior Manager, Infrastructure Platform Engineering to lead a team building core systems that turn large... ...AKS) Familiarity with the operational challenges of GPU clusters, AI training, and inference workloads Working knowledge of platform security...TrainingTemporary workImmediate start$190k - $215k
...customers running production AI inference workloads and... ...Inference, model serving platforms, LLM deployments, AI agents, or GPU‑based inference environments... ...with the product and engineering teams to define customer... ...infrastructure trends. Customer Training: Deliver training...TrainingFull timeTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Platform Engineer: GPU Inference & Training. Be the first to apply!
- data platform engineer San Francisco, CA
- client platform engineer San Francisco, CA
- platform engineer San Francisco, CA
- platform developer San Francisco, CA
- senior platform engineer San Francisco, CA
- platform engineering manager San Francisco, CA
- director of digital platform San Francisco, CA
- digital platform specialist San Francisco, CA
- power platform San Francisco, CA
- platform manager San Francisco, CA

