Member of Technical Staff - RL Training Framework
PVH (Tommy Hilfiger/Calvin Klein)
The RL infrastructure team is looking for an engineer to help develop our RL training framework Design and implement the systems backing all RL workloads at xAI, from small scale ablations to production training runs Profile, debug, and optimize end-to-end training performance Improve scalability and observability of the RL stack Benefits Health and wellness: Comprehensive health insurance including medical, dental, vision, and disability coverage Life and family: Life and AD&D insurance and fertility benefits to ensure our team’s well-being and peace of mind Flexible vacation: We work hard but avoid burn out. Take time off when you need it Visa sponsorship: We support international talent with visa sponsorship to join our team 401(k) plan: Retirement savings plan to secure your financial future Experience building, debugging, and optimizing efficiency of large-scale distributed systems. Comfortable diving into unfamiliar areas and solving problems at all levels of the stack. Proficiency in Python, Jax, Rust, and/or C++. Experience with large scale LLM training infrastructure. Strong knowledge of reinforcement learning techniques. Experience with RL numerics. #J-18808-Ljbffr PVH (Tommy Hilfiger/Calvin Klein)
$180k
...concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement the systems backing all RL workloads...TrainingFull timeTemporary work$180k
...able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement the systems backing all RL workloads at SpaceXAI,...TrainingTemporary work- ...applied AI research lab pioneering data and RL environment curation for training and evaluating agents. Recently, we... ...rollouts and identifying subtle patterns. Technical execution: Proficiency in Python and ML frameworks (PyTorch, JAX, or similar). Experience with...TrainingFlexible hours
- About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll... ...of LLM reasoning, post-training, and agentic system design.... ...training (instruction tuning, RL, reasoning) Data pipeline... ...familiarity with modern ML frameworks (e.g., PyTorch, JAX, or TensorFlow...Training
- Member of Technical Staff — Kernel / Compiler / Communication About the Role RadixArk is seeking a Member... .... This role is critical to scaling training and inference across thousands of... ...and developed Miles, our large‑scale RL framework. We build world‑class systems for AI...TrainingFlexible hours
- RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference... ...and developed Miles (our large‑scale RL framework). We're on a mission to democratize... ...class open systems for inference and training. Our team has optimized kernels...TrainingWorldwideFlexible hours
- About The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that... ...inference optimization for large-scale RL or other production workload.... ...etc.) Familiarity with post-training framework (e.g. Miles, Slime, veRL, Prime-RL,...TrainingFlexible hours
- RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier... ...running training jobs Develop training frameworks and infrastructure tooling... ...and developed Miles, our large-scale RL framework. We build world-class infrastructure...TrainingFlexible hours
$180k
...teammates. ABOUT THE ROLE: You will work on the most critical post‑training and reinforcement learning challenges at any given time — including reward modeling, preference optimization (RLHF/DPO), and RL for improving reasoning, truthfulness, and real‑world capabilities...TrainingTemporary work- ...intersection of data, evaluation, and model training. We design evaluation harnesses... ...-designing datasets, evaluation frameworks, and training and RL pipelines that shape how their... .... About The Role We are hiring Members of Technical Staff to build the data and evaluation...TrainingShift work
- ...What You’ll Do As a Founding Member of the Technical Staff at Architect, you'll be at the forefront of training AI models for chip design,... ...where you'll own the end-to-end RL workflow—from reward... ...computing, and distributed training frameworks (e.g., PyTorch, CUDA, QLoRA,...Training
- RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps SGLang... ...and developed Miles (our large-scale RL framework). We’re on a mission to democratize... ...class open systems for inference and training. Our team has optimized kernels...TrainingFlexible hoursNight shift
- About The Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems... ...with distributed inference systems (SGLang, vLLM) or training/RL frameworks (Miles, Megatron, veRL, TorchTitan) CPU inference...TrainingFlexible hours
- About The Role RadixArk is seeking a Member of Technical Staff — Diffusion Model to advance the frontier... ...—from designing novel algorithms to training and deploying models at scale. Your... ...and developed Miles (our large‑scale RL framework). We’re on a mission to democratize...TrainingFlexible hours
- Member of Technical Staff — Supercomputing About the Role RadixArk is hiring a Member of Technical Staff... ...for frontier‑scale inference and training workloads. This role sits at the intersection... ...developed Miles , our large‑scale RL framework. We are building world‑class open...TrainingFlexible hours
$180k
...About the Role The mid‑training team at xAI aims to provide... ...boost the ceiling for RL. Engineer long‑context... ...Spark, Ray, and other frameworks for large‑scale data... ...interview”) during which a member of our team will ask... ...which consists of four technical interviews: Coding assessment...TrainingTemporary workRelocation$180k - $250k
Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk... ...-performance inference and training systems using JAX, XLA, and... ...JAX, XLA, or TPU-focused frameworks Bachelor's or Master's... ...developed Miles (our large-scale RL framework). We're on a...TrainingFull timeFlexible hours$180k
Member of Technical Staff - Multimodal Understanding About xAI xAI’s mission is... ...curation/acquisition, tokenizer training, large‑scale pre‑training,... ...models. Create evaluation frameworks, internal benchmarks,... ...stack (pre‑training > SFT/RL/post‑training) to enable reasoning...TrainingTemporary work$180k
...reinforcement learning data for training Grok reasoning models. Our... ...training data, and advancing RL algorithms. About the Role... ...phone interview”) during which a member of our team will ask some basic... ...process, which consists of four technical interviews: # Coding...TrainingTemporary workWork at officeWork from homeRelocation- Member of Technical Staff — Developer Technology About the Role RadixArk is seeking... ...to make LLM inference and training dramatically faster,... ...reinforcement-learning post-training framework for large-scale LLM and MoE... ...silicon Training systems: RL post‑training with Miles,...TrainingFlexible hours
- Member of Technical Staff - Developer Experience About the Role RadixArk is seeking a Developer Advocate... ..., including LLM serving, distributed training, and GPU optimization. Proficiency... ...and developed Miles (our large-scale RL framework). We're on a mission to democratize...TrainingFlexible hours
$180k
...through synthetic data generation. Optimize mid-training data mixtures to boost the ceiling for RL. Engineer long-context data recipes.... ...Strong engineering abilities in Spark, Ray, and other frameworks for large-scale data processing. COMPENSATION AND...TrainingFull timeTemporary work- About the Role As a Member of Technical Staff [Platform] at NeoCognition , you’ll design and build the... ...developer tools, automation frameworks, or internal platforms . Familiarity... ...with machine learning infrastructure , training pipelines, or model evaluation tooling...Training
- ...world models as simulators, feature extractors, and training grounds for real robot policies and model‑based RL approaches. Most new experiments here will fail;... ...ownership of the full ML stack, including core frameworks that Odyssey researchers and product engineers rely...TrainingRemote workFlexible hours
- ...Reinforcement Learning experiments (GRPO/PPO/DPO), training data mixes and reward signal explorations... ...are also encouraged to apply. RL Knowledge: Strong academic understanding... ...proficiency in Python and deep learning frameworks (PyTorch). You should be able to write clean...TrainingInternship
- ...who can push LLM inference and training systems to the limit across... ...workloads, cost-per-million tokens, RL rollout efficiency, and... ...and cloud partners on deep technical evaluations Contribute performance... ...Miles, our large-scale RL framework. We're on a mission to...TrainingFlexible hours
$148.5k - $223.9k
Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member of Technical Staff... ...Hands-on experience with deep learning frameworks** *Strong understanding of ML... ...Experience implementing and debugging model training, evaluation, and inference pipelines*...TrainingWork at office$200k - $300k
...This team works on foundational model capabilities, post-training techniques, building RL infra and infrastructure that benefits the entire... ...model development Build robust and effective training frameworks (on top of Megatron/PyTorch) for post-training LLMs Implement...TrainingFull time$180k
...Responsibilities span data curation, modeling, training, inference serving, and product... ...long‑horizon synthesis, agentic planning, RL training, and world simulation (including... ...visual and audio data. Design evaluation frameworks, metrics, benchmarks, evals, and reward models...TrainingTemporary work- ...real robot tasks. Post-training at Rhoda means taking a... ...levels — from senior to staff. What You'll Do Design and implement RL training pipelines to improve... ...Build evaluation frameworks for post-trained policies... ...are expected to define technical direction and drive research...TrainingShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff - RL Training Framework. Be the first to apply!
- work from home technical support specialist Palo Alto, CA
- product support technician Palo Alto, CA
- helpdesk support technician Palo Alto, CA
- help desk assistant Palo Alto, CA
- IT help desk technician Palo Alto, CA
- technical solutions specialist Palo Alto, CA
- desktop support analyst Palo Alto, CA
- trade support analyst Palo Alto, CA
- technical analyst Palo Alto, CA
- technical support specialist Palo Alto, CA

