Research Member of Technical Staff- Training Systems
Rhoda AI
At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $400M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality. We're looking for a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end. You will define how our models train at scale — driving efficiency, scalability, and correctness across large-scale multimodal training. This is a core systems role, not infrastructure support. Your work directly determines how efficiently we use compute, how well models scale across thousands of GPUs, and how quickly research can iterate. What You'll Do Own training performance end-to-end Diagnose and improve performance of large-scale multimodal training (vision, video, proprioception, actions, language) Build systematic performance attribution: step-time decomposition (compute vs communication vs input pipeline), scaling curves across cluster sizes, and bottleneck identification and prioritization Drive measurable gains in: Distributed efficiency (comm/compute overlap, bucketization, topology-aware mapping, parallelism strategies) Compute efficiency (kernel hotspots, operator fusion, attention optimization, framework/runtime overhead) Memory efficiency (activation checkpointing, sequence packing/bucketing, fragmentation reduction) Design training systems (not just tune them) Define and evolve parallelism strategies: data / tensor / pipeline / sharding / hybrid approaches Improve execution efficiency through communication scheduling and overlap, graph capture and execution optimization, and runtime-level improvements Contribute to and extend training frameworks where needed Make performance observable and measurable Establish source-of-truth performance metrics: step-time breakdowns, MFU / throughput / scaling efficiency Build tools to identify bottlenecks quickly, track performance across model families, and compare scaling behavior across configurations Develop regression detection: microbenchmarks, performance baselines, and automated detection of efficiency regressions Partner deeply with researchers Work side-by-side with research scientists and research engineers — no silos Translate model innovations into scalable, efficient implementations Advise on training tradeoffs for robotics world models: long-horizon sequences, rollout/evaluation cadence, multimodal and variable-length data Collaborate on cluster-level efficiency Work with infrastructure/SRE teams to improve utilization across large distributed jobs, impact of network and collective performance on training, and topology-aware job placement and scaling behavior What We're Looking For Proven track record improving large-scale distributed training performance Deep hands-on experience with modern ML stacks (PyTorch required; JAX a plus) Strong understanding of data / tensor / pipeline parallelism, sharded training (FSDP / ZeRO-style), communication patterns and overlap strategies, and scaling behavior across large GPU clusters Strong systems intuition — ability to reason across compute, communication, and memory bottlenecks Exceptional debugging and measurement ability: turn “training is slow” into clear bottlenecks, experiments, and validated improvements High ownership mindset and comfort in a fast-moving environment Nice to Have (But Not Required) GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion) Experience with multimodal or video training (variable-length sequences, packing/bucketing) Experience working on large-scale training frameworks or distributed runtimes Familiarity with cluster topology, networking, and large-scale scheduling effects Why This Role Direct leverage on research velocity — every efficiency gain you make accelerates model iteration across the entire research team Own the scalability and performance of large-scale multimodal training for real-world embodied intelligence, not static benchmarks Improvements you make compound across every training run the company executes — high ownership, high impact, small elite team #J-18808-Ljbffr Rhoda AI
- ...from high-performance hardware and robot systems to the infrastructure and state-of-the-... ...cases, made possible by our cutting edge research and end-to-end system design. We've... ...Research Engineer to build and maintain the training platform that powers our model development...Technical training
- ...performance hardware and robot systems to the infrastructure... ...by our cutting edge research and end-to-end system... ...robot tasks. Post-training at Rhoda means taking... ...levels — from senior to staff. What You'll Do... ...are expected to define technical direction and drive research...Technical trainingShift work
- ...performance hardware and robot systems to the infrastructure and... ...possible by our cutting edge research and end-to-end system design.... ...We're looking for a senior or staff-level Research Engineer or ML... ...through dataset generation, model training, inference, and real-robot...Suggested
- About The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large-scale... ...infrastructure tooling Collaborate with model researchers to support frontier experiments Debug and...Technical trainingFlexible hours
- RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models. You will work on large-scale distributed... ...infrastructure tooling Collaborate with model researchers to support frontier experiments Debug and resolve...Technical trainingFlexible hours
$180k
Member of Technical Staff - Pre-Training About xAI xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization...Technical trainingTemporary work$180k
...xAI’s mission is to create AI systems that can accurately... ...teammates. About the Role The mid‑training team at xAI aims to provide an... ...phone interview”) during which a member of our team will ask some... ...process, which consists of four technical interviews: Coding...Technical trainingTemporary workRelocation$180k
Member of Technical Staff, Pre-training Data Infrastructure xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. All employees...Technical trainingTemporary workRelocation- Member of Technical Staff, Applied AI Research San Francisco, CA; Sunnyvale, CA DoorDash’s mission is to empower local... ...customers, to make our logistics system even more efficient. We have an... ...replace - human decision‑making; trained personnel make final decisions with...Hourly payWork at officeLocal areaImmediate startFlexible hours
- ...paradigms. Born out of Stanford Research, our team blends AI with... ...What You’ll Do As a Founding Member of the Technical Staff at Architect, you'll be at the forefront of training AI models for chip design,... ...and structured coding tasks. Systems Engineering: Strong software...
$180k - $250k
Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems... ...high-performance inference and training systems using JAX, XLA, and Pallas.... ...that powers leading AI companies and research labs. Join us in building...Full timeFlexible hours$148.5k - $223.9k
Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member of Technical Staff - AI Research page is loaded## Senior Member of Technical... ...and iterate agentic AI systems with customers. With... ...implementing and debugging model training, evaluation, and...Work at office- About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll be part of the core team advancing... ...the frontier of LLM agents — systems that can reason, plan, and act... ...in the areas of LLM reasoning, post-training, and agentic system design. Develop...
$180k
...SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.... ...flash models through synthetic data generation. Optimize mid-training data mixtures to boost the ceiling for RL. Engineer long-...Technical trainingFull timeTemporary work$180k
...focused on what comes next: agentic systems that interact naturally with people... ...the Role We are looking for a Member of Technical Staff - Mid-Training to lead the development of training... ...and PyTorch; comfort working across research and systems code. - Ability to work...Technical trainingFull time$180k
...SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.... ...infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement...Technical trainingFull timeTemporary work- ...analyze, and optimize hardware systems, integrating cutting-edge... ...Role Summary Hands-on technical role spanning validation, training data, and customers.... ...bringing what you learn back to research and product as feature... ...for the right person — members of the team already work...Technical trainingFull timeRemote work2 days per week3 days per week
$180k
...Description Job Description SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in... .... You are a power user of AI models. If you previously trained models used by millions of people it's a big plus, but modeling...Technical trainingTemporary work$180k
SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our... ...team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement the systems backing...Technical trainingTemporary work- The RL infrastructure team is looking for an engineer to help develop our RL training framework Design and implement the systems backing all RL workloads at xAI, from small scale ablations to production training runs Profile, debug, and optimize end-to-end training performance...Technical trainingVisa sponsorshipFlexible hours
- Job You will own the training pipeline behind the models that power both Parallel’s search stack and Parallel’s agents... ...models that serve all three. You care about your research being applied to product and systems that millions use. Compensation & benefits Competitive...Technical trainingWork at officeVisa sponsorship
$175k - $350k
.... About the Role As a Model Training engineer, you will design, build... ...on the fun parts. Balance research curiosity with product... ...Communicate crisply with both technical and non-technical teammates.... ...improvements in customer-facing systems. Salary Range : $175,000 - $...Technical training$148.5k - $223.9k
...future of Salesforce. Salesforce AI Research is looking for a Machine Learning... ...implement and iterate agentic AI systems with customers. With your strong technical competence, strategic thinking... ...track records, such as LLM, pre/post‑training, RL, agentic system. Prioritizes...- About Architect Architect is an AI research and product lab for chip design. We build AI models and systems that can explore, design, optimize, and verify new hardware... ...Learning experiments (GRPO/PPO/DPO), training data mixes and reward signal explorations. Contribute...Internship
$200k - $420k
...personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research. Who we are We are scientists, engineers, and... ...a proven track record of scaling consumer systems for hundreds of millions of users and architecting...Local areaVisa sponsorshipRelocation package$148.5k - $223.9k
...engineering, product, and AI-focused activities across agentic AI systems. Specific responsibilities are not enumerated as a separate... ...frameworks and strong ML fundamentals; experience debugging model training, evaluation, and inference pipelines. Infrastructure &...- Parallel Web Systems in Palo Alto is looking for a skilled engineer to build and manage their large-scale full-text indexing systems.... ...distributed systems. Join us in our fully in-person team committed to solving complex technical challenges. #J-18808-Ljbffr Parallel Web Systems
$200k - $300k
...Department AI Perplexity is seeking top-tier AI Research Scientists and Engineers to advance our... ...foundational model capabilities, post-training techniques, building RL infra and... ...with large-scale LLMs and Deep Learning systems Strong programming skills in Python/PyTorch...Full time- ...Role DoorDash is building an AI Research org from the ground up, and... ...High compute budgets for training and inference, sized to support... ...environments built on real operational systems, training and evaluation... ...shapes how our team members move quickly, learn, and reiterate...Hourly payWork at officeLocal areaRemote workFlexible hours
- ...robot task performance Research and implement data... ...Collaborate closely with pre-training and post-training teams... ..., and actionable Staff-level candidates are expected to define technical direction and drive research... ...of large-scale systems and generative model research...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Member of Technical Staff- Training Systems. Be the first to apply!
- art history research assistant Mountain View, CA
- biochemistry research assistant Mountain View, CA
- research fellow dermatology Mountain View, CA
- research associate in biology Mountain View, CA
- research administrator Mountain View, CA
- social science research assistant Mountain View, CA
- research assistant Mountain View, CA
- research associate scientist Mountain View, CA
- computer science research assistant Mountain View, CA
- neuroscience research assistant Mountain View, CA

