Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Member of Technical Staff- Training Systems

Rhoda AI

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $400M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality. We're looking for a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end. You will define how our models train at scale — driving efficiency, scalability, and correctness across large-scale multimodal training. This is a core systems role, not infrastructure support. Your work directly determines how efficiently we use compute, how well models scale across thousands of GPUs, and how quickly research can iterate. What You'll Do Own training performance end-to-end Diagnose and improve performance of large-scale multimodal training (vision, video, proprioception, actions, language) Build systematic performance attribution: step-time decomposition (compute vs communication vs input pipeline), scaling curves across cluster sizes, and bottleneck identification and prioritization Drive measurable gains in: Distributed efficiency (comm/compute overlap, bucketization, topology-aware mapping, parallelism strategies) Compute efficiency (kernel hotspots, operator fusion, attention optimization, framework/runtime overhead) Memory efficiency (activation checkpointing, sequence packing/bucketing, fragmentation reduction) Design training systems (not just tune them) Define and evolve parallelism strategies: data / tensor / pipeline / sharding / hybrid approaches Improve execution efficiency through communication scheduling and overlap, graph capture and execution optimization, and runtime-level improvements Contribute to and extend training frameworks where needed Make performance observable and measurable Establish source-of-truth performance metrics: step-time breakdowns, MFU / throughput / scaling efficiency Build tools to identify bottlenecks quickly, track performance across model families, and compare scaling behavior across configurations Develop regression detection: microbenchmarks, performance baselines, and automated detection of efficiency regressions Partner deeply with researchers Work side-by-side with research scientists and research engineers — no silos Translate model innovations into scalable, efficient implementations Advise on training tradeoffs for robotics world models: long-horizon sequences, rollout/evaluation cadence, multimodal and variable-length data Collaborate on cluster-level efficiency Work with infrastructure/SRE teams to improve utilization across large distributed jobs, impact of network and collective performance on training, and topology-aware job placement and scaling behavior What We're Looking For Proven track record improving large-scale distributed training performance Deep hands-on experience with modern ML stacks (PyTorch required; JAX a plus) Strong understanding of data / tensor / pipeline parallelism, sharded training (FSDP / ZeRO-style), communication patterns and overlap strategies, and scaling behavior across large GPU clusters Strong systems intuition — ability to reason across compute, communication, and memory bottlenecks Exceptional debugging and measurement ability: turn “training is slow” into clear bottlenecks, experiments, and validated improvements High ownership mindset and comfort in a fast-moving environment Nice to Have (But Not Required) GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion) Experience with multimodal or video training (variable-length sequences, packing/bucketing) Experience working on large-scale training frameworks or distributed runtimes Familiarity with cluster topology, networking, and large-scale scheduling effects Why This Role Direct leverage on research velocity — every efficiency gain you make accelerates model iteration across the entire research team Own the scalability and performance of large-scale multimodal training for real-world embodied intelligence, not static benchmarks Improvements you make compound across every training run the company executes — high ownership, high impact, small elite team #J-18808-Ljbffr Rhoda AI

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Research Member of Technical Staff- Training Systems in Mountain View, CA vacancy
  •  ...from high-performance hardware and robot systems to the infrastructure and state-of-the-...  ...cases, made possible by our cutting edge research and end-to-end system design. We've...  ...Research Engineer to build and maintain the training platform that powers our model development... 
    Technical training

    Rhoda AI

    Mountain View, CA
    10 hours ago
  •  ...performance hardware and robot systems to the infrastructure...  ...by our cutting edge research and end-to-end system...  ...robot tasks. Post-training at Rhoda means taking...  ...levels — from senior to staff. What You'll Do...  ...are expected to define technical direction and drive research... 
    Technical training
    Shift work

    Rhoda AI

    Mountain View, CA
    10 hours ago
  •  ...performance hardware and robot systems to the infrastructure and...  ...possible by our cutting edge research and end-to-end system design....  ...We're looking for a senior or staff-level Research Engineer or ML...  ...through dataset generation, model training, inference, and real-robot... 
    Suggested

    Rhoda AI

    Mountain View, CA
    3 days ago
  • About The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large-scale...  ...infrastructure tooling Collaborate with model researchers to support frontier experiments Debug and... 
    Technical training
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large-scale distributed...  ...infrastructure tooling Collaborate with model researchers to support frontier experiments Debug and resolve... 
    Technical training
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • $180k

    Member of Technical Staff - Pre-Training About xAI xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization... 
    Technical training
    Temporary work

    Xai

    Palo Alto, CA
    10 hours ago
  • $180k

     ...xAI’s mission is to create AI systems that can accurately...  ...teammates. About the Role The mid‑training team at xAI aims to provide an...  ...phone interview”) during which a member of our team will ask some...  ...process, which consists of four technical interviews: Coding... 
    Technical training
    Temporary work
    Relocation

    Pantera Capital

    Palo Alto, CA
    2 days ago
  • $180k

    Member of Technical Staff, Pre-training Data Infrastructure xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. All employees... 
    Technical training
    Temporary work
    Relocation

    xAI

    Palo Alto, CA
    1 day ago
  • Member of Technical Staff, Applied AI Research San Francisco, CA; Sunnyvale, CA DoorDash’s mission is to empower local...  ...customers, to make our logistics system even more efficient. We have an...  ...replace - human decision‑making; trained personnel make final decisions with... 
    Hourly pay
    Work at office
    Local area
    Immediate start
    Flexible hours

    DoorDash USA

    Sunnyvale, CA
    4 days ago
  •  ...paradigms. Born out of Stanford Research, our team blends AI with...  ...What You’ll Do As a Founding Member of the Technical Staff at Architect, you'll be at the forefront of training AI models for chip design,...  ...and structured coding tasks. Systems Engineering: Strong software... 

    Architect Labs

    Palo Alto, CA
    1 day ago
  • $180k - $250k

    Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems...  ...high-performance inference and training systems using JAX, XLA, and Pallas....  ...that powers leading AI companies and research labs. Join us in building... 
    Full time
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago
  • $148.5k - $223.9k

    Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member of Technical Staff - AI Research page is loaded## Senior Member of Technical...  ...and iterate agentic AI systems with customers. With...  ...implementing and debugging model training, evaluation, and... 
    Work at office

    Salesforce, Inc.

    Palo Alto, CA
    10 hours ago
  • About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll be part of the core team advancing...  ...the frontier of LLM agents — systems that can reason, plan, and act...  ...in the areas of LLM reasoning, post-training, and agentic system design. Develop... 

    NeoCognition Inc.

    Palo Alto, CA
    4 days ago
  • $180k

     ...SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge....  ...flash models through synthetic data generation. Optimize mid-training data mixtures to boost the ceiling for RL. Engineer long-... 
    Technical training
    Full time
    Temporary work

    SpaceXAI

    Palo Alto, CA
    2 days ago
  • $180k

     ...focused on what comes next: agentic systems that interact naturally with people...  ...the Role We are looking for a Member of Technical Staff - Mid-Training to lead the development of training...  ...and PyTorch; comfort working across research and systems code. - Ability to work... 
    Technical training
    Full time

    Hark

    San Jose, CA
    2 days ago
  • $180k

     ...SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge....  ...infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement... 
    Technical training
    Full time
    Temporary work

    SpaceXAI

    Palo Alto, CA
    2 days ago
  •  ...analyze, and optimize hardware systems, integrating cutting-edge...  ...Role Summary Hands-on technical role spanning validation, training data, and customers....  ...bringing what you learn back to research and product as feature...  ...for the right person — members of the team already work... 
    Technical training
    Full time
    Remote work
    2 days per week
    3 days per week

    Vinci4d

    Palo Alto, CA
    2 days ago
  • $180k

     ...Description Job Description SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in...  .... You are a power user of AI models. If you previously trained models used by millions of people it's a big plus, but modeling... 
    Technical training
    Temporary work

    SpaceXAI

    Palo Alto, CA
    a month ago
  • $180k

    SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our...  ...team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement the systems backing... 
    Technical training
    Temporary work

    Neura Market

    Palo Alto, CA
    3 days ago
  • The RL infrastructure team is looking for an engineer to help develop our RL training framework Design and implement the systems backing all RL workloads at xAI, from small scale ablations to production training runs Profile, debug, and optimize end-to-end training performance... 
    Technical training
    Visa sponsorship
    Flexible hours

    PVH (Tommy Hilfiger/Calvin Klein)

    Palo Alto, CA
    3 days ago
  • Job You will own the training pipeline behind the models that power both Parallel’s search stack and Parallel’s agents...  ...models that serve all three. You care about your research being applied to product and systems that millions use. Compensation & benefits Competitive... 
    Technical training
    Work at office
    Visa sponsorship

    Parallel Web Systems

    Palo Alto, CA
    4 days ago
  • $175k - $350k

     .... About the Role As a Model Training engineer, you will design, build...  ...on the fun parts. Balance research curiosity with product...  ...Communicate crisply with both technical and non-technical teammates....  ...improvements in customer-facing systems. Salary Range : $175,000 - $... 
    Technical training

    Inflection AI

    Palo Alto, CA
    2 days ago
  • $148.5k - $223.9k

     ...future of Salesforce. Salesforce AI Research is looking for a Machine Learning...  ...implement and iterate agentic AI systems with customers. With your strong technical competence, strategic thinking...  ...track records, such as LLM, pre/post‑training, RL, agentic system. Prioritizes... 

    salesforce.com, inc.

    Palo Alto, CA
    4 days ago
  • About Architect Architect is an AI research and product lab for chip design. We build AI models and systems that can explore, design, optimize, and verify new hardware...  ...Learning experiments (GRPO/PPO/DPO), training data mixes and reward signal explorations. Contribute... 
    Internship

    Architect Labs

    Palo Alto, CA
    1 day ago
  • $200k - $420k

     ...personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research. Who we are We are scientists, engineers, and...  ...a proven track record of scaling consumer systems for hundreds of millions of users and architecting... 
    Local area
    Visa sponsorship
    Relocation package

    River AI Inc.

    Palo Alto, CA
    4 days ago
  • $148.5k - $223.9k

     ...engineering, product, and AI-focused activities across agentic AI systems. Specific responsibilities are not enumerated as a separate...  ...frameworks and strong ML fundamentals; experience debugging model training, evaluation, and inference pipelines. Infrastructure &... 

    Salesforce

    Palo Alto, CA
    2 days ago
  • Parallel Web Systems in Palo Alto is looking for a skilled engineer to build and manage their large-scale full-text indexing systems....  ...distributed systems. Join us in our fully in-person team committed to solving complex technical challenges. #J-18808-Ljbffr Parallel Web Systems

    Parallel Web Systems

    Palo Alto, CA
    4 days ago
  • $200k - $300k

     ...Department AI Perplexity is seeking top-tier AI Research Scientists and Engineers to advance our...  ...foundational model capabilities, post-training techniques, building RL infra and...  ...with large-scale LLMs and Deep Learning systems Strong programming skills in Python/PyTorch... 
    Full time

    Pantera Capital

    Palo Alto, CA
    10 hours ago
  •  ...Role DoorDash is building an AI Research org from the ground up, and...  ...High compute budgets for training and inference, sized to support...  ...environments built on real operational systems, training and evaluation...  ...shapes how our team members move quickly, learn, and reiterate... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash

    Sunnyvale, CA
    10 hours ago
  •  ...robot task performance Research and implement data...  ...Collaborate closely with pre-training and post-training teams...  ..., and actionable Staff-level candidates are expected to define technical direction and drive research...  ...of large-scale systems and generative model research... 

    Rhoda AI

    Mountain View, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Member of Technical Staff- Training Systems. Be the first to apply!