Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC

the company serves hundreds of millions of queries a month, and every one of them fans out into multiple AI inference requests running in real time. Behind that sits a large GPU fleet spread across several cloud providers. Today, our inference engineers and researchers build models while also managing networking, securing capacity, and operating the underlying GPU clusters, responsibilities we want a dedicated platform team to own. Your job is to take ownership of that infrastructure and hide its complexity behind a unified, self-serve platform for running training and inference workloads. Responsibilities Build a self-serve compute platform. Design and own the systems that let inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure. Operate the GPU fleet. Own provisioning, lifecycle management, reliability, and capacity integration across providers, giving teams a consistent way to use compute regardless of where it runs. Solve for GPU scarcity. Build the scheduling and placement logic that finds available capacity across providers, packs it efficiently, and gets the right workload onto the right hardware under real constraints. Support two very different workloads. Keep long-running distributed training jobs healthy while simultaneously guaranteeing the availability and latency of production inference services on the same fleet. Own the Kubernetes for GPU orchestration. Write the operators and CRDs, and manage many clusters across providers so the platform behaves the same everywhere we run. Make failure boring. Build the fault tolerance, autoscaling, and observability that keep the fleet utilized and let workloads survive node loss, provider hiccups, and capacity shifts without human intervention. Set technical direction across teams. Partner with inference and cloud infrastructure engineers to turn operational constraints into a coherent platform architecture and roadmap. Qualifications Deep Kubernetes experience — custom operators, CRDs, and multi-cluster federation, not just running kubectl apply. You've managed GPU clusters at scale: NVIDIA hardware, CUDA, and the networking that makes them fast (InfiniBand or RoCE). You've orchestrated compute across multiple clouds (CoreWeave, AWS, GCP, or similar) and understand how different each one really is. Strong distributed systems fundamentals: scheduling, resource allocation, and fault tolerance under load. You write infrastructure and systems-level code in Go, Rust or C++. You've supported both long-running training jobs and high-availability inference services, and you know why they pull infrastructure in opposite directions. You own problems end-to-end and do well when the path forward isn't laid out for you. Additional experience we value Inference serving stacks: vLLM, SGLang, or TensorRT-LLM. Slurm or other HPC schedulers. GPU kernel work in CUDA or Triton — not required, but notable. High-speed interconnects: InfiniBand, RoCE, or RDMA in production. Observability for ML workloads: Prometheus, Grafana, or Weights & Biases. If you’re excited about this role, we encourage you to apply even if your experience doesn’t match every qualification listed above. #J-18808-Ljbffr United States Digital Space LLC

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff (Software Engineer, Inference & Training Platform) in San Francisco, CA vacancy
  • $150k - $300k

     ...anyone to create, train, and deploy them....  ...LLM serving, LLM inference optimization and RL...  ...training stack. Core Technical Responsibilities...  ...LLM serving platform that operates across...  ...PyTorch: LLM Inference engine development and...  ...and encourage team members to contribute to the... 
    Platform
    Training
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...serve humanity. We’re training and deploying...  ...team of researchers, engineers, designers, and more...  ...next generation of AI platforms powering advanced NLP...  ...We are looking for Members of Technical Staff to join the Model Serving...  ...and throughput of inference. ~ Strong... 
    Platform
    Training
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    15 hours ago
  •  ...development at scale. Our platform combines powerful distributed training infrastructure with an...  ...enabling researchers and engineers to train state‑of‑the‑...  ...training systems Core Technical Responsibilities Platform...  ...development and encourage team members to contribute to the... 
    Platform
    Training
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Kubelt

    San Francisco, CA
    4 days ago
  •  ...it's simple to serve low-latency inference, fine-tune models, and access...  ...olympiad medalists, and experienced engineering and product leaders with decades...  ...serving time. We're building a platform that covers the whole life of an LLM -- train it, deploy it, observe it -- and... 
    Platform
    Training
    Work at office

    Mixpeek

    San Francisco, CA
    1 day ago
  • $190.9k - $232.8k

    P-1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development...  ...cross-team collaboration: with platform engineers, cloud infrastructure, and...  ..., relevant certifications and training, and specific work location. Based... 
    Platform
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $190k - $265k

     ...and AI infrastructure platform so our customers can use...  ...business. Founded by engineers — and customer-...  ...opportunity to solve technical challenges, from designing...  ...apps, AI agents, model training, model serving, and Vector...  ...real-time and batch inference, powering model inference... 
    Platform
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  •  ...Perplexity is seeking energetic engineers to join our highly driven Agents engineering team...  ...include Perplexity Computer (our platform for generalized frontier intelligence),...  ...valuable units of work for our users; Training action and decision models that determine... 
    Platform
    Training
    Full time
    Flexible hours

    Perplexity

    San Francisco, CA
    2 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply...  ...founding member of the engineering team, you will impact...  ...with our data-centric platform designed to simplify and...  ..., transformation, training/fine-tuning, and inference? You will also: Find opportunities... 
    Platform
    Training
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    3 days ago
  •  ...experiences. As a growth engineer at Perplexity, you'll...  ...-to-end projects from training and productionizing...  ...years of professional software engineering experience...  ...testing and experimentation platforms (Eppo, Statsig,...  ...into simple, effective technical solutions. Self-motivated... 
    Platform
    Training

    aijoblist

    San Francisco, CA
    3 days ago
  •  ...for fast, efficient inference. As AI workloads...  ...together. Gimlet's platform intelligently...  ...than headcount. The engineers we hire today will...  ...will be built on software capable of intelligently...  ...of education, training, and professional...  ...company. As an early member of the team, you... 
    Platform
    Training

    The Consensus

    San Francisco, CA
    1 day ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is...  ...frontier-scale pre-training, domain-specific RL...  ...long context, and inference-time compute to...  ...About the role As an engineer on the...  ...looking for Strong software engineering skills... 
    Platform
    Training
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    1 day ago
  • Careers / Member of Technical Staff (AI research) Member of Technical...  ...core parts of Fearn’s platform from deploying our custom...  ...play a crucial role in training models and designing inference pipelines, pushing the...  ..., particularly Compute Engine, GKE, Vertex AI, or Cloud... 
    Platform
    Training
    Full time
    Work at office

    Kindredventures

    San Francisco, CA
    15 hours ago
  • Member of Technical Staff, Applied AI The opportunity We are...  ...learners, protein engineers and biologists, jointly...  ...protein screening platforms. At Latent Labs...  ...architectures, training dynamics and inference behaviour. You are...  ...production enterprise software. You have... 
    Platform
    Training
    Flexible hours

    Latent Labs

    San Francisco, CA
    4 days ago
  • $150k - $300k

     ...lets anyone create, train, and deploy them....  ...hosted training platform - the product that...  ...runs the jobs. Core Technical Responsibilities...  ...based training and inference orchestration...  ...We're looking for engineers who are fluent across...  ...and encourage team members to contribute to the... 
    Platform
    Training
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Kubelt

    San Francisco, CA
    4 days ago
  •  ...intelligence. As a Member of Technical Staff, you'll be at...  ...collaborating with platform teams to ensure...  ...efficient model inference, video...  ...closely with robotics engineers to integrate your...  ...world datasets to train and deploy state...  ...in professional software development... 
    Platform
    Training
    Local area

    Amazon Science

    San Francisco, CA
    1 day ago
  • $200k - $400k

     ...for F100 enterprises, trained a first-of-its-kind...  ...best researchers, engineers, designers, and operators...  ...means running inference over populations of...  ...About the Role As a Member of Technical Staff in Research Infrastructure...  ...you will build the platform our researchers... 
    Platform
    Training
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    3 days ago
  •  .... Our Emerald Conductor software platform makes data centers flexible...  .... About the Role As a Member of Technical Staff, you will help invent...  ...with hands‑on software engineering experience and are excited...  ...Kubernetes, Slurm, distributed training/inference frameworks, or large‑... 
    Platform
    Training
    Work from home
    Flexible hours
    2 days per week

    Emerald AI

    San Francisco, CA
    1 day ago
  • $250k

    Eragon — Member of Technical Staff Type: Full-time | On-site | San...  ...operating system. It post-trains open-source models...  ...Systems engineering: Design scalable pipelines...  ...pipelines for training, inference, and data processing...  ...page also shows a platform "Pending Approval" stage... 
    Platform
    Training
    Full time
    H1b
    Work at office
    Local area
    Visa sponsorship

    davidjoseph-co

    San Francisco, CA
    3 days ago
  •  ...underlying hardware. Our platform intelligently...  ...Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will...  ...This role is ideal for engineers who deeply understand...  ...Qualifications Strong software engineering fundamentals... 
    Platform

    Gimlet Labs

    San Francisco, CA
    1 day ago
  •  ...company. We support labs, engineers and enterprises to understand...  ...Language model inference is the fastest-moving market...  ...point, and we’re hiring a Member of Technical Staff to drive them. You’ll own...  ...our inference benchmarking platform together with our engineers... 
    Platform
    Shift work

    Artificial Analysis

    San Francisco, CA
    2 days ago
  • Member of Technical Staff - Agents at Prime Intellect - San Francisco...  ...code or capital to train powerful, open...  ...research, and other engineering teams to identify key...  ...Requirements Agent & Platform Skills Python (FastAPI...  ...optimize agent training or inference on GPUs. Advanced AI... 
    Platform
    Training
    Remote work
    Flexible hours

    Victrays

    San Francisco, CA
    5 hours ago
  •  ...AI-powered answer engine built to serve the...  ...worker that can use software like a human to...  ...the Role The Data Platform team owns the end-...  .... In this senior/staff role, you will shape...  ...the long-term technical direction of Perplexity...  ...features, AI training and evaluation workflows... 
    Platform
    Training

    Perplexity

    San Francisco, CA
    3 days ago
  • $250k - $300k

     ...in the market, building a platform that deploys and operates large...  ...next-generation AI training and inference at scale. This role offers...  ...platform evolves while solving engineering challenges that have a...  ...track record of impressive technical work you can speak to in depth... 
    Platform
    Training
    Full time
    Remote work
    San Francisco, CA
    9 days ago
  • Member of Technical Staff, Lead Researcher San Francisco, CA; Sunnyvale, CA...  ...and mentor researchers, engineers, and fellows joining the...  ...across DoorDash with ML platform, product, and operations...  ...High compute budgets for training and inference, sized to support frontier... 
    Platform
    Training
    Local area

    DoorDash USA

    San Francisco, CA
    3 days ago
  • $220k - $250k

     ...patient care, is hiring a Member of Technical Staff to join their team...  ...working closely with the core platform from day one while helping...  ...Create robust evaluation and training frameworks, including...  ...systems. Skillset ~Strong software engineering skills with a focus on... 
    Platform
    Training
    Full time
    Remote work

    Alldus International Consulting Ltd

    San Francisco, CA
    more than 2 months ago
  • Member of Technical Staff - Software Engineer Valthos | Posted Mar 3 Full-time Negotiable Advanced (5-10 yrs) Software Engineer Valthos Inc. Valthos...  ...stack software engineers to design and implement a platform that will allow Valthos to bring frontier methods to... 
    Platform
    Full time
    Work at office

    Valthos

    San Francisco, CA
    1 day ago
  •  ...infrastructure platform for...  ...that enable ML engineers and data scientists...  ...spans ML training...  ...serve as a technical lead across...  ...products. As a Staff Engineer, you...  ...latency model inference, large-scale...  ...operating great software systems.Who...  ...mentor team members.Experience building... 
    Platform
    Training
    Flexible hours

    Stripe

    San Francisco, CA
    3 days ago
  • Responsibilities Design, deploy, and maintain large distributed ML training and inference clusters Develop efficient, scalable end-to-end pipelines...  ...) to train large foundation models Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings... 
    Platform
    Training

    Kindredventures

    San Francisco, CA
    15 hours ago
  • $150k - $300k

     ...enables anyone to create, train, and deploy them. We...  ...build the developer platform that makes all of it...  ...This is a generalist software engineering role focused on...  ...experience improvements Technical Requirements Strong...  ...development and encourage team members to contribute to the... 
    Platform
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...GPU clusters that power training, evaluation, and serving...  ...tenancy across training and inference workloads Build software that abstracts cluster...  ...interface to researchers and engineers Own cluster storage and...  ...code Knowledge of cloud platforms including GCP, AWS, or... 
    Platform
    Training

    Linuxcareers

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff (Software Engineer, Inference & Training Platform). Be the first to apply!