Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff, GPU / ML Systems

Full-time

SkyPilot

Role Description

SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on:

  • Accelerator scheduling and utilization
  • The serving path
  • The integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads

A few points of GPU utilization here can save a team millions in compute and days on every training run.

What you'll do

  • Own GPU scheduling, utilization and health:
    • How SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes
    • Real-time GPU health monitoring and automatic failure recovery
  • Build optimizations for training and serving:
    • Enable large scale pre-training with node hot-swapping
    • Design storage systems for fast model checkpointing
    • Container migration, inference autoscaling and multi-cluster serving
    • Preemption handling and sandboxes for training, RL rollouts, and evals
  • Make the AI stack run great out of the box:
    • Deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference

Qualifications

  • Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure
  • Strong Python, and comfort reaching into systems-level and GPU-adjacent details
  • You care about squeezing most from the compute available to you
  • Experience operating large-scale training or high-throughput inference in production

Requirements

  • Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe)
  • You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after

Benefits

  • Competitive compensation and equity
  • Comprehensive medical, dental, vision coverage for you and your dependents
  • The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership
  • A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale)
  • Gourmet lunch & dinner for the team to do their best work

Company Description

SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."

SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff, GPU / ML Systems in Remote vacancy
  • $200k - $400k

     ...Description We're looking for a TPU and AMD GPU performance engineer to make vLLM a first-...  .... ~Work at the boundary of inference systems, kernels, compilers, and hardware architecture...  ...workloads. ~Ability to optimize ML inference paths such as attention, GEMM, sampling... 
    Suggested
    Full time
    Remote work
    Visa sponsorship

    Inferact

    Remote
    a month ago
  • Role Description We're looking for an AMD GPU performance engineer to make vLLM a first-...  ...You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture...  ...constraints. ~Experience optimizing ML kernels or inference paths such as attention... 
    Suggested
    Full time

    Inferact

    Remote
    a month ago
  • $300k

     ...a full-duplex audiovisual system that can listen, speak, react...  ...We’re looking for a deeply technical Member of Technical Staff to own RL and post‑training...  ..., serving latency, GPU utilization, policy update...  ...training, alignment, evaluation, ML systems, or model behavior.... 
    Suggested
    H1b
    Work at office
    Visa sponsorship
    Shift work

    Nuance Labs, Inc.

    Seattle, WA
    5 days ago
  •  ...sophisticated infrastructure and software systems. We are seeking a Senior Member of Technical Staff to build core systems, solve...  ...skills. Requirements ~AI/ML infrastructure. ~Kubernetes and...  ...Startup or scale-up experience. ~GPU or data-intensive systems.... 
    Suggested
    Full time

    Pragmatike

    Remote
    15 days ago
  • Member of the Technical Staff, Systems Location: North America Remote / San Francisco, CA Full-Time About Andromeda Compute is the most sought-after...  ...tens of thousands of GPUs under management and billions of GPU-hours supported, on everything from A100 to GB300. The... 
    Suggested
    Full time
    Remote work

    Andromeda

    San Francisco, CA
    2 days ago
  • Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member...  ...data scientists and ML engineers to focus on building...  ...compute resources (CPU and GPU) efficiently? What data model...  ...related field 5+ years of systems engineering experience in an... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    5 days ago
  • $230k

    Cerebras Systems builds the world's largest AI chip, 56 times larger...  ...run large-scale ML applications, without the hassle...  ..., over 10 times faster than GPU-based hyperscale cloud inference...  ...has multiple openings for Sr. Member of Technical Staff. Title: Sr. Member of... 
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $150k - $300k

     ...that powers our training systems Our developer-facing...  ...reliable at scale. Core Technical Responsibilities...  ...heterogeneous hardware (CPU, GPU, TPU) Platform...  ...with GPU computing and ML infrastructure Knowledge...  ...development and encourage team members to contribute to the... 
    Full time
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Menlo Ventures

    San Francisco, CA
    3 days ago
  •  ...enterprises who are building AI systems. We believe that our work is...  ..., and join the team. As a member of technical staff with a focus on multimodal...  ...in writing efficient GPU kernels using CUDA, optimising...  ...comfortable diving into complex ML codebases to identify and resolve... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco...  ...behavior, traces, workload replay, GPU signals, and task-path...  ...ambiguity, learn new systems quickly, and turn rough customer...  ...Experience with inference or ML infrastructure.* Exposure to... 
    Full time
    Remote work

    Touchdown Labs, Inc.

    San Francisco, CA
    1 day ago
  • $150k - $300k

     ...inference optimization and RL systems. You will be working on...  ...training stack. Core Technical Responsibilities LLM...  ...across our cloud GPU fleets. GPU‑Aware Scheduling...  ...Experience Building ML Systems at Scale: 3+...  ...development and encourage team members to contribute to the... 
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    2 days ago
  •  ...Description Job Description Member of Technical Staff, Machine Learning,...  ...Learning, you will build core ML components.  The Member of Technical...  ...work on real production systems from day one, learning how...  ...improvement. - Tech Stack:  GPU, JAX, ML, Machine Learning,... 
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    12 days ago
  •  ...sophisticated infrastructure and software systems. We are seeking a Staff level Member of Technical Staff to lead high-impact...  ...authority. Requirements ~AI/ML infrastructure (Nice to Have)....  ...up experience (Nice to Have). ~GPU or data-intensive systems (Nice to... 
    Full time

    Pragmatike

    Remote
    15 days ago
  • Member of Technical Staff, Model Efficiency Who are we? Our mission is to scale intelligence...  ...who are building AI systems to power magical experiences...  ...on building reliable ML systems and pushing the boundaries...  ...techniques, including GPU/CUDA optimizations, kernel-level... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  • Member of Technical Staff - Agents at Prime Intellect - San Francisco Building the Future of Open Source...  ...Deployments : Architect and maintain systems that support distributed AI agent...  ...effective infrastructure. Nice to Have GPU/ML Infrastructure : Understanding how to... 
    Remote work
    Flexible hours

    Victrays

    San Francisco, CA
    4 days ago
  • $150k - $300k

     ...our success (not the inverse). A team member summarized: “Roboflow is a company...  ...What You'll Do As a Member of Technical Staff on our Frontier Data team, you’ll build...  ...autonomy across a wide surface area — systems engineering, ML tooling, and research collaboration —... 
    Full time
    Remote work
    Work from home
    Home office
    Relocation package

    Roboflow

    Remote
    4 days ago
  • $95.2k - $165.5k

     ...hardest problems and providing unmatched technical expertise. As the operator of a...  ...satellite, launch, ground, and cyber systems for defense, civil and commercial customers...  ...a Communications Systems Engineer (Member of Technical Staff/Senior Member of Technical Staff -... 
    Full time
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    aero

    El Segundo, CA
    4 days ago
  • Member of Technical Staff - Engineer, RL and Control Position Summary You will help...  ...—domain randomization, system identification, and the iterative...  ...). Experience with ML experiment-tracking and training...  ...the-art machine learning and GPU programming frameworks, such... 
    Work from home
    Flexible hours

    Walden Robotics, Inc.

    Cambridge, MA
    3 days ago
  •  ...Consulting Member Of Technical Staff As a Consulting Member of Technical Staff, you will be a key contributor...  ...creation of a robust and intelligent system, enhancing the healthcare experience...  ...~ Hands-on experience building AI/ML or generative AI applications, including... 
    Immediate start
    Remote work
    Worldwide

    Hackajob

    United States
    4 days ago
  • $200k - $300k

     ...Member of Technical Staff — ML Infra (Data) Seattle, Washington About Nuance Labs Nuance Labs is building photorealistic, real-time AI avatars with emotional intelligence: a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real... 
    H1b
    Work at office
    Visa sponsorship

    Nuance Labs, Inc.

    Seattle, WA
    5 days ago
  • $180k - $280k

     ...repetitive workflows, and data that no existing system handles well - and we're building the AI to fix it. As a Member of the Technical Staff at Finch, you'll own the full lifecycle of...  ...accuracy and reliability. Own the full ML lifecycle - model selection, fine-tuning, prompt... 
    Work at office
    Remote work
    Flexible hours
    1 day per week

    Finch Services

    New York, NY
    5 days ago
  •  ...Perimeter Institute building AI systems to discover new physics at...  ...observability and alerting, schedule GPU and CPU capacity, and lead...  ...pipelines, observability for ML systems. Comprehensive operational...  ...‑stage equity. We evaluate on technical breadth, systems thinking,... 
    Remote work

    Physical Superintelligence

    Boston, MA
    3 days ago
  • $150k - $300k

    Role Description As a Member of Technical Staff on our Frontier Data team, you’ll build the environments, evaluations, and datasets that...  ...operate with a lot of autonomy across a wide surface area — systems engineering, ML tooling, and research collaboration — and your judgment... 
    Full time
    Work from home
    Home office

    Roboflow

    Remote
    3 days ago
  •  ...partner and the regulators depend on. The system serves hundreds of accounts today. The...  ...It's the product. Why the title is Member of Technical Staff Because they'd rather hire the person...  ...re not filtering on them. # An AI or ML background. We run models against live... 
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours

    EQL Tech

    San Francisco, CA
    3 days ago
  • Role Description We are looking for a Member of Technical Staff, Research to investigate, design, test...  ...computer science, data science, information systems, or related field. ~Proven track...  ...development processes, in-depth AI/ML concepts, and data infrastructure.... 
    Full time
    Remote work

    FirstPrinciples

    Remote
    a month ago
  •  ...practical challenges of making AI systems work at scale. This role demands deep technical expertise in production...  ...can do your best work. As a Member of Technical Staff, you will: ~Design and write...  ...Proficiency in Python and related ML frameworks such as JAX,... 
    Full time
    Local area
    Remote work

    Cohere

    Remote
    16 days ago
  • $220k - $250k

     ...personalizing patient care, is hiring a Member of Technical Staff to join their team remotely. The...  ...help build and scale the AI and data systems powering next-generation primary care...  ...designing and deploying production-grade AI/ML systems. ~Solid understanding of PostgreSQL... 
    Full time
    Remote work

    Alldus International Consulting Ltd

    San Francisco, CA
    more than 2 months ago
  •  ...edge research into production-ready machine learning systems ~ Design, build, and deploy end-to-end ML models and pipelines ~ Develop and optimize models...  .... Our founding team, along with many of our team members, has contributed to many of the breakthroughs in AI... 
    Full time
    H1b
    Remote work
    Visa sponsorship

    Reka

    Remote
    13 days ago
  • $200k - $320k

     ...Member of Technical Staff, Robotics Research Job Type: Full-time   Location: Remote   The...  ...generation robot learning and perception systems. You’ll collaborate with leading AI...  ...learning, embodied AI, or applied ML research. Strong understanding of robotics... 
    Full time
    Temporary work
    Remote work

    ESR Healthcare

    New York, NY
    11 days ago
  • $250k - $300k

     ...they have built a platform that deploys GPU clusters into third-party datacentres at...  ...designing and building high-performance systems spanning storage, networking, virtualisation...  ...Must Have: A track record of impressive technical work you can speak to in depth; the years... 
    Full time
    Remote work
    San Francisco, CA
    23 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff, GPU / ML Systems. Be the first to apply!