Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff - RL Training Framework

PVH (Tommy Hilfiger/Calvin Klein)

The RL infrastructure team is looking for an engineer to help develop our RL training framework Design and implement the systems backing all RL workloads at xAI, from small scale ablations to production training runs Profile, debug, and optimize end-to-end training performance Improve scalability and observability of the RL stack Benefits Health and wellness: Comprehensive health insurance including medical, dental, vision, and disability coverage Life and family: Life and AD&D insurance and fertility benefits to ensure our team’s well-being and peace of mind Flexible vacation: We work hard but avoid burn out. Take time off when you need it Visa sponsorship: We support international talent with visa sponsorship to join our team 401(k) plan: Retirement savings plan to secure your financial future Experience building, debugging, and optimizing efficiency of large-scale distributed systems. Comfortable diving into unfamiliar areas and solving problems at all levels of the stack. Proficiency in Python, Jax, Rust, and/or C++. Experience with large scale LLM training infrastructure. Strong knowledge of reinforcement learning techniques. Experience with RL numerics. #J-18808-Ljbffr PVH (Tommy Hilfiger/Calvin Klein)

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff - RL Training Framework in Palo Alto, CA vacancy
  • $180k

     ...concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement the systems backing all RL workloads... 
    Training
    Full time
    Temporary work

    SpaceXAI

    Palo Alto, CA
    2 days ago
  • $180k

     ...able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The RL infrastructure team is looking for an engineer to help develop our RL training framework. RESPONSIBILITIES: Design and implement the systems backing all RL workloads at SpaceXAI,... 
    Training
    Temporary work

    Neura Market

    Palo Alto, CA
    3 days ago
  •  ...applied AI research lab pioneering data and RL environment curation for training and evaluating agents. Recently, we...  ...rollouts and identifying subtle patterns. Technical execution: Proficiency in Python and ML frameworks (PyTorch, JAX, or similar). Experience with... 
    Training
    Flexible hours

    BespokeLabs.AI, Inc

    Mountain View, CA
    2 days ago
  • About the Role As a Member of Technical Staff [Research] at NeoCognition , you’ll...  ...of LLM reasoning, post-training, and agentic system design....  ...training (instruction tuning, RL, reasoning) Data pipeline...  ...familiarity with modern ML frameworks (e.g., PyTorch, JAX, or TensorFlow... 
    Training

    NeoCognition Inc.

    Palo Alto, CA
    4 days ago
  • Member of Technical Staff — Kernel / Compiler / Communication About the Role RadixArk is seeking a Member...  .... This role is critical to scaling training and inference across thousands of...  ...and developed Miles, our large‑scale RL framework. We build world‑class systems for AI... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    2 days ago
  • RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference...  ...and developed Miles (our large‑scale RL framework). We're on a mission to democratize...  ...class open systems for inference and training. Our team has optimized kernels... 
    Training
    Worldwide
    Flexible hours

    Dormont Manufacturing Co

    Palo Alto, CA
    4 days ago
  • About The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that...  ...inference optimization for large-scale RL or other production workload....  ...etc.) Familiarity with post-training framework (e.g. Miles, Slime, veRL, Prime-RL,... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier...  ...running training jobs Develop training frameworks and infrastructure tooling...  ...and developed Miles, our large-scale RL framework. We build world-class infrastructure... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • $180k

     ...teammates. ABOUT THE ROLE: You will work on the most critical post‑training and reinforcement learning challenges at any given time — including reward modeling, preference optimization (RLHF/DPO), and RL for improving reasoning, truthfulness, and real‑world capabilities... 
    Training
    Temporary work

    Pantera Capital

    Palo Alto, CA
    9 hours ago
  •  ...intersection of data, evaluation, and model training. We design evaluation harnesses...  ...-designing datasets, evaluation frameworks, and training and RL pipelines that shape how their...  .... About The Role We are hiring Members of Technical Staff to build the data and evaluation... 
    Training
    Shift work

    Orbifold AI

    Palo Alto, CA
    1 day ago
  •  ...What You’ll Do As a Founding Member of the Technical Staff at Architect, you'll be at the forefront of training AI models for chip design,...  ...where you'll own the end-to-end RL workflow—from reward...  ...computing, and distributed training frameworks (e.g., PyTorch, CUDA, QLoRA,... 
    Training

    Architect Labs

    Palo Alto, CA
    1 day ago
  • RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps SGLang...  ...and developed Miles (our large-scale RL framework). We’re on a mission to democratize...  ...class open systems for inference and training. Our team has optimized kernels... 
    Training
    Flexible hours
    Night shift

    RadixArk

    Palo Alto, CA
    3 days ago
  • About The Role RadixArk is seeking a Member of Technical Staff: Accelerator Systems to push the limits of performance for frontier AI systems...  ...with distributed inference systems (SGLang, vLLM) or training/RL frameworks (Miles, Megatron, veRL, TorchTitan) CPU inference... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • About The Role RadixArk is seeking a Member of Technical Staff — Diffusion Model to advance the frontier...  ...—from designing novel algorithms to training and deploying models at scale. Your...  ...and developed Miles (our large‑scale RL framework). We’re on a mission to democratize... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    9 hours ago
  • Member of Technical Staff — Supercomputing About the Role RadixArk is hiring a Member of Technical Staff...  ...for frontier‑scale inference and training workloads. This role sits at the intersection...  ...developed Miles , our large‑scale RL framework. We are building world‑class open... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • $180k

     ...About the Role The mid‑training team at xAI aims to provide...  ...boost the ceiling for RL. Engineer long‑context...  ...Spark, Ray, and other frameworks for large‑scale data...  ...interview”) during which a member of our team will ask...  ...which consists of four technical interviews: Coding assessment... 
    Training
    Temporary work
    Relocation

    Pantera Capital

    Palo Alto, CA
    2 days ago
  • $180k - $250k

    Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk...  ...-performance inference and training systems using JAX, XLA, and...  ...JAX, XLA, or TPU-focused frameworks Bachelor's or Master's...  ...developed Miles (our large-scale RL framework). We're on a... 
    Training
    Full time
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago
  • $180k

    Member of Technical Staff - Multimodal Understanding About xAI xAI’s mission is...  ...curation/acquisition, tokenizer training, large‑scale pre‑training,...  ...models. Create evaluation frameworks, internal benchmarks,...  ...stack (pre‑training > SFT/RL/post‑training) to enable reasoning... 
    Training
    Temporary work

    xAI

    Palo Alto, CA
    9 hours ago
  • $180k

     ...reinforcement learning data for training Grok reasoning models. Our...  ...training data, and advancing RL algorithms. About the Role...  ...phone interview”) during which a member of our team will ask some basic...  ...process, which consists of four technical interviews: # Coding... 
    Training
    Temporary work
    Work at office
    Work from home
    Relocation

    xAI

    Palo Alto, CA
    more than 2 months ago
  • Member of Technical Staff — Developer Technology About the Role RadixArk is seeking...  ...to make LLM inference and training dramatically faster,...  ...reinforcement-learning post-training framework for large-scale LLM and MoE...  ...silicon Training systems: RL post‑training with Miles,... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago
  • Member of Technical Staff - Developer Experience About the Role RadixArk is seeking a Developer Advocate...  ..., including LLM serving, distributed training, and GPU optimization. Proficiency...  ...and developed Miles (our large-scale RL framework). We're on a mission to democratize... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago
  • $180k

     ...through synthetic data generation. Optimize mid-training data mixtures to boost the ceiling for RL. Engineer long-context data recipes....  ...Strong engineering abilities in Spark, Ray, and other frameworks for large-scale data processing. COMPENSATION AND... 
    Training
    Full time
    Temporary work

    SpaceXAI

    Palo Alto, CA
    2 days ago
  • About the Role As a Member of Technical Staff [Platform] at NeoCognition , you’ll design and build the...  ...developer tools, automation frameworks, or internal platforms . Familiarity...  ...with machine learning infrastructure , training pipelines, or model evaluation tooling... 
    Training

    NeoCognition

    Palo Alto, CA
    4 days ago
  •  ...world models as simulators, feature extractors, and training grounds for real robot policies and model‑based RL approaches. Most new experiments here will fail;...  ...ownership of the full ML stack, including core frameworks that Odyssey researchers and product engineers rely... 
    Training
    Remote work
    Flexible hours

    Odyssey

    Palo Alto, CA
    3 days ago
  •  ...Reinforcement Learning experiments (GRPO/PPO/DPO), training data mixes and reward signal explorations...  ...are also encouraged to apply. RL Knowledge: Strong academic understanding...  ...proficiency in Python and deep learning frameworks (PyTorch). You should be able to write clean... 
    Training
    Internship

    Architect Labs

    Palo Alto, CA
    1 day ago
  •  ...who can push LLM inference and training systems to the limit across...  ...workloads, cost-per-million tokens, RL rollout efficiency, and...  ...and cloud partners on deep technical evaluations Contribute performance...  ...Miles, our large-scale RL framework. We're on a mission to... 
    Training
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  • $148.5k - $223.9k

    Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member of Technical Staff...  ...Hands-on experience with deep learning frameworks** *Strong understanding of ML...  ...Experience implementing and debugging model training, evaluation, and inference pipelines*... 
    Training
    Work at office

    Salesforce, Inc.

    Palo Alto, CA
    9 hours ago
  • $200k - $300k

     ...This team works on foundational model capabilities, post-training techniques, building RL infra and infrastructure that benefits the entire...  ...model development Build robust and effective training frameworks (on top of Megatron/PyTorch) for post-training LLMs Implement... 
    Training
    Full time

    Pantera Capital

    Palo Alto, CA
    9 hours ago
  • $180k

     ...Responsibilities span data curation, modeling, training, inference serving, and product...  ...long‑horizon synthesis, agentic planning, RL training, and world simulation (including...  ...visual and audio data. Design evaluation frameworks, metrics, benchmarks, evals, and reward models... 
    Training
    Temporary work

    xAI

    Palo Alto, CA
    3 days ago
  •  ...real robot tasks. Post-training at Rhoda means taking a...  ...levels — from senior to staff. What You'll Do Design and implement RL training pipelines to improve...  ...Build evaluation frameworks for post-trained policies...  ...are expected to define technical direction and drive research... 
    Training
    Shift work

    Rhoda AI

    Mountain View, CA
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff - RL Training Framework. Be the first to apply!