Member of Technical Staff, GPU / ML Systems
SkyPilot
Role Description
SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on:
- Accelerator scheduling and utilization
- The serving path
- The integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads
A few points of GPU utilization here can save a team millions in compute and days on every training run.
What you'll do
- Own GPU scheduling, utilization and health:
- How SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes
- Real-time GPU health monitoring and automatic failure recovery
- Build optimizations for training and serving:
- Enable large scale pre-training with node hot-swapping
- Design storage systems for fast model checkpointing
- Container migration, inference autoscaling and multi-cluster serving
- Preemption handling and sandboxes for training, RL rollouts, and evals
- Make the AI stack run great out of the box:
- Deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference
Qualifications
- Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure
- Strong Python, and comfort reaching into systems-level and GPU-adjacent details
- You care about squeezing most from the compute available to you
- Experience operating large-scale training or high-throughput inference in production
Requirements
- Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe)
- You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after
Benefits
- Competitive compensation and equity
- Comprehensive medical, dental, vision coverage for you and your dependents
- The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership
- A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale)
- Gourmet lunch & dinner for the team to do their best work
Company Description
SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."
SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.
$200k - $400k
...Description We're looking for a TPU and AMD GPU performance engineer to make vLLM a first-... .... ~Work at the boundary of inference systems, kernels, compilers, and hardware architecture... ...workloads. ~Ability to optimize ML inference paths such as attention, GEMM, sampling...SuggestedFull timeRemote workVisa sponsorship$200k - $400k
Role Description We're looking for an AMD GPU performance engineer to make vLLM a first-... ...You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture... ...constraints. ~Experience optimizing ML kernels or inference paths such as attention...SuggestedFull time$160k - $190k
Role Description As a Member of Technical Staff, Machine Learning, you will build core ML components. The Member of Technical Staff will work on real production systems from day one, learning how large-scale... ...improvement. ~Tech Stack: GPU, JAX, ML, Machine Learning,...SuggestedFull timeRemote work$250k
...platform. ~Design and own the systems that let inference engineers... ...services without managing GPU provisioning, cluster configuration... ...human intervention. ~Set technical direction across teams: ~... ...production. ~Observability for ML workloads: Prometheus, Grafana...SuggestedFull timeShift work- ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member... ...data scientists and ML engineers to focus on building... ...compute resources (CPU and GPU) efficiently? What data model... ...a related field 5+ years of systems engineering experience in an...SuggestedFull timePart timeWork at officeWork from homeFlexible hours2 days per week
$300k
...a full-duplex audiovisual system that can listen, speak, react... ...We’re looking for a deeply technical Member of Technical Staff to own RL and post‑training... ..., serving latency, GPU utilization, policy update... ...training, alignment, evaluation, ML systems, or model behavior....H1bWork at officeVisa sponsorshipShift work- Role Description We are looking for a Member of Technical Staff, Research to investigate, design, test... ...computer science, data science, information systems, or related field. ~Proven track... ...development processes, in-depth AI/ML concepts, and data infrastructure....Full timeRemote work
- ...designing and building high performance systems across, but not limited to: ~Storage... ...runtimes Qualifications ~Impressive technical work you can go deep on, with impact in the... ...of GPUs under management. ~Billions of GPU-hours supported, on everything from A100...Full time
- ...machine learning, generative AI, agent-based systems, and graph technologies to get our... ...products. We are looking for a Founding Member of Technical Staff to join our team guiding the development and deployment of complex ML systems reporting to Dan Wald, Co-Founder...Work at officeRemote work
$250k
...understanding of what dangers AI systems pose, and we are extremely... ...unsuccessful execution in frontier ML research. You are undaunted by... ...systems and make sound technical decisions. You lead large projects... ...in the growth of your team members. You can hold a team to high standards...H1bWork at officeWork from homeHome officeRelocation packageFlexible hours3 days per week- ...Member of Technical Staff, Machine Learning Protingent Staffing has an exciting direct hire Member of Technical Staff,... ...Technical Staff, Machine Learning, you will build core ML components. You will work on real production systems from day one, learning how large‑scale ML...Remote work
$200k - $300k
...Member of Technical Staff — ML Infra (Data) Seattle, Washington About Nuance Labs Nuance Labs is building photorealistic, real-time AI avatars with emotional intelligence: a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real...H1bWork at officeVisa sponsorship$180k - $280k
...Join the Flock Careers at finch Member of Technical Staff, Machine Learning Location New York City Employment... ...workflows, and data that no existing system handles well — and we’re building the... ...accuracy and reliability. Own the full ML lifecycle — model selection, fine‑...Full timeWork at officeRemote workFlexible hours1 day per week- ...that optimizes AI itself. Our journey starts with GPU kernels, but will expand into every corner of ML systems and AI infrastructure. We're a small team (4... ...Look For You're a strong fit if you: Have deep technical intuition and can learn new domains quickly Are...Remote work
- ...-world task completion. The system must handle multi-step reasoning... .... Role As a Senior Member of Technical Staff, Machine Learning, you are... ...independent owner of critical ML subsystems in production.... ...Python PyTorch / JAX GPU-based training and inference...
$200k - $400k
...We're looking for an inference systems educator-builder: someone who... ...inference systems project, teach hard technical concepts clearly, and create... ...decode ~Quantization ~GPU serving ~Latency versus... ...Requirements ~Prior work in ML systems, distributed systems,...Full time- ...ONNX Runtime and DeepSpeed with deep expertise in distributed systems and large-scale model training infrastructure Varun is an... .... You’ll work closely with our founding designer and backend / ML teams to ship new product surfaces, workflows, and interactions...Full time
- ...edge research into production-ready machine learning systems ~Design, build, and deploy end-to-end ML models and pipelines ~Develop and optimize models... .... Our founding team, along with many of our team members, has contributed to many of the breakthroughs in AI...Full timeH1bRemote work
- ...engine that turns every investigation into a compounding asset: the more the platform runs, the sharper it gets. This is applied ML systems, not research. If your goal is training foundation models from scratch and publishing, this is not the seat. Location: Remote (...Full timeRemote work
$200k - $300.09k
...Description The OpenClaw Foundation is seeking exceptional Members of Technical Staff (MTS) to serve as full-time maintainers, builders, and... ...developers, researchers, enterprises, and end users. ~Build systems that support large-scale AI agent deployment and...Full time- ...re looking for an engineer to own this core end to end, set its technical direction and build new features to make it more robust,... ...world. What you'll do: ~Design and build the future of AI systems: solve some of the hardest problems in distributed AI systems to...Full time
- ...Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking engineers to... ...environments. You ship; you do not stall. Fluency at the physics‑ML intersection: you do not need to be an ML expert, but you can read...Remote work
$324k - $396k
...SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge... ...and accurately share knowledge with their teammates. Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques...Remote work- ...partner and the regulators depend on. The system serves hundreds of accounts today. The... ...It's the product. Why the title is Member of Technical Staff Because they'd rather hire the person... ...re not filtering on them. # An AI or ML background. We run models against live...Work at officeRemote workVisa sponsorshipFlexible hours
- ...performance workloads, and you'll design the cloud compute, distributed systems, and sandboxed tooling that keeps them reliable, efficient, and... ...-source contributions or production experience with large-scale ML infrastructure Compensation, benefits, and perks We offer...Work at officeRemote workFlexible hours
$220k - $250k
...personalizing patient care, is hiring a Member of Technical Staff to join their team remotely. The... ...help build and scale the AI and data systems powering next-generation primary care... ...designing and deploying production-grade AI/ML systems. ~Solid understanding of PostgreSQL...Permanent employmentRemote work- ...everyone. Patients interact with advanced AI systems at every step of their care and follow... .... About the role We are hiring Members of Technical Staff (MTS) across a range of seniority... ...AI in healthcare. As an MTS on the AI/ML team, you will design, build, and ship...Full timeRemote workFlexible hours
$140k - $300k
.... About the Role We’re hiring a Member of Technical Staff to operate as a builder–operator at the... ...building new things over maintaining systems You think like a founder: ownership... ...technical background (software, systems, ML, or similar) Experience building...Full time$120k - $150k
Role Description Join our team as a Member of Technical Staff, AI/ML Engineering, and drive the development of cutting-edge AI systems. Work alongside talented engineers to design, implement, and optimize advanced LLM-based solutions, leveraging your expertise in quantitative...Full timeContract workRemote work$250k - $300k
...they have built a platform that deploys GPU clusters into third-party datacentres at... ...designing and building high-performance systems spanning storage, networking, virtualisation... ...Must Have: A track record of impressive technical work you can speak to in depth; the years...Permanent employmentRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff, GPU / ML Systems. Be the first to apply!
- technical assistant Remote
- work from home technical support specialist Remote
- end user support technician Remote
- help desk technical support Remote
- technical support assistant Remote
- tech assistant Remote
- remote support technician Remote
- application support technician Remote
- technical associate Remote
- technical analyst Remote















