GPU Systems Engineer — Distributed Training & Inference
TensorScale AI
TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical improvements that scale. Responsibilities include CUDA / Triton optimizations, designing efficient distributed inference and training pipelines, and owning communication performance across GPUs and nodes with #J-18808-Ljbffr TensorScale AI
- ...in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional...Training
- ...excellence. We seek engineers with strong intrinsic... ...We’re looking for a systems engineer with HPC or... ...experience to help scale AI inference. You’ll leverage your... ...systems to optimize GPU performance at the... ...Familiarity with distributed training/inference frameworks...TrainingFull timeWork at office
- ...of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution... ...will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-...Training
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...Suggested
- ...San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will... ...ensure efficient hardware utilization across multi‑node GPU clusters. Join a team focused on scalable AI foundations...Training
$200.8k - $251k
...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary...TrainingFull time- ...Baseten powers mission-critical inference for the world's most... ...and help build the platform engineers turn to to ship AI products... ...building the global operating system for distributed, heterogeneous AI hardware.... ...engineers to lead our GPU Networking efforts, making...Full timeFlexible hours
$250k
...building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company... ...Site Reliability Engineer to support and scale... ...observability across distributed compute environments... ...available infrastructure systems Improve CI/CD...TrainingFull timeRemote work- ...Francisco is seeking a Staff ML Systems Engineer to design and prototype... ...low-latency, high-throughput inference. You will implement changes... ...systems, while profiling across GPU, networking, and memory to... ...also co-design RL and post-training pipelines, drive performance...Training
- ...AI Systems Engineer - Codex Core Agents Location San Francisco... ..., model behavior, inference/runtime issues, and product... ...systems in distributed systems, infrastructure... ...model evals, or post‑training feedback loops. Background... ...optimization, GPU systems, benchmarking...TrainingFull timeWork at officeLocal areaRelocation packageFlexible hours
- ...BeamBeam is an ultrafast AI inference platform. We built a... ...runtime that launches GPU-backed containers in... ...help us with Platform Engineering work. We're working on... ...problems:Low-level systems development: working with... ...working with a large distributed systemComfortable with...
$300 per month
...seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s Hardware... ...across Crusoe Cloud’s GPU- and CPU-based infrastructure... ...studies across training and inference - dense, MoE, long-context... ....Hands-on experience with distributed training and/or inference...TrainingTemporary work$350k
...and steerable AI systems. We want AI to be... ...committed researchers, engineers, policy experts,... ...Anthropic's inference fleet serves Claude... ...kernels, model servers, distributed routing,... ...systems, especially training or inference infrastructure... ...Familiarity with GPU/TPU/accelerator...TrainingFull timeWork at officeVisa sponsorshipFlexible hours- ...foundation Model and seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute environment. You will... ...provisioning, upgrading, and optimizing distributed clusters that power training and inference workloads. You will extend orchestration,...Training
$225k
...Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and... ...high reliability for RL and post-training workflows. The ideal candidate will possess...Training- ...design, build, and operate GPU-heavy infrastructure for high-throughput model inference and mid-training workloads. Join a team that... ...generation, RL pipelines, and distributed model evaluation across thousands... ..., and scalable distributed systems. #J-18808-Ljbffr Visa HuntTraining
$150k - $300k
...anyone to create, train, and deploy them.... ...Solutions Architect for GPU Infrastructure,... ...production‑ready systems capable of... ...for LLM training, inference, and HPC workloads... ...and Kubernetes for distributed workloads Implement... ...with our world‑class engineering team while having...Training- ...in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible... ...ideal candidate has a strong background in distributed systems and is eager to engage in...
- ...is recruiting infrastructure engineers to scale large-scale inference and evaluation around a... ...will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes... ...efficient inference stacks, GPU-aware optimization, and deep...
$227.2k - $417k
...the Role:As a Software Engineer on the ML... ...class machine learning inference platforms. These platforms... ...latency ML model serving systems that support Deep Learning... ...throughput, and low latency distributed systems using... ...), ElastiCache, model training orchestration, etc.Understanding...TrainingFull timeTemporary workLocal areaFlexible hours- ...’re hiring a hands‑on Vision Systems Engineer to own the detection, tracking... ...Discrimination: Develop and train lightweight CNN classifiers for... ...for real‑time embedded inference via quantization, pruning, or... ...detect — on embedded FPGA and GPU platforms Performance Modeling...Training
$150k - $300k
...enables anyone to create, train, and deploy them. We... ...LLM serving, LLM inference optimization and RL systems. You will be working... ...across our cloud GPU fleets. GPU‑Aware Scheduling... ...SLOs. Model Distribution: Optimize model... ...PyTorch: LLM Inference engine development and integration...TrainingWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- ...an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting... ...will work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding...Training
- ...technologies—like the da Vinci surgical system and Ion—have transformed how... ...worldwide.We’re a team of engineers, clinicians, and innovators... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you... ...models to real-time onboard inference—while serving as a core...Local areaWorldwideFlexible hours
$160k - $194k
...organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of... ...level, skill sets, market conditions, experience and training, licensure and certifications, and business and...TrainingFull timeRemote workFlexible hours- ...Fluidstack Production Engineering TeamExamples of key exciting problems... ...actual state inspection, and distributed command execution. One... ...not a hundred scripts.Make the system's view of itself always match... ...platforms, so every new site and GPU generation lands cleanly from...Local area
$110 per hour
.... Position: MLOps Engineer (JAX, PyTorch, Pallas/... ...performance in MLOps , training infrastructure, and... ...to MLOps and ML systems problems . Evaluate... ...training pipeline design, distributed systems reasoning, and... ...or optimizing custom GPU kernels using Pallas...TrainingRemote jobContract workSummer workWeekday work- ...company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on...Training
$117.2k - $223.9k
...experiences. Join our team of talented engineers and help us advance the... ...infrastructure security, and the security of distributed and scalable distributed systems. This role requires hands-on... ...compensation, promotion, benefits, training, assessment of job performance, discipline...TrainingFull time- ...practical constraints of robotic platforms. About the Role As a Research Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage distributed data pipelines,...TrainingFull timeWork at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Systems Engineer — Distributed Training & Inference. Be the first to apply!
- senior windows systems engineer San Francisco, CA
- software system engineer San Francisco, CA
- system test engineer San Francisco, CA
- mission system engineer San Francisco, CA
- healthcare systems engineer San Francisco, CA
- electronic systems engineer San Francisco, CA
- operating system engineer San Francisco, CA
- system engineer remote San Francisco, CA
- application system engineer San Francisco, CA
- advanced systems engineer San Francisco, CA




