Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

RL Infrastructure Engineer Scalable Training & Performance

Xai

xAI in Palo Alto is seeking an experienced engineer to design and implement the RL training framework and the systems backing all RL workloads, from ablations to production runs. You will profile, debug, and optimize end-to-end training performance, and improve scalability and observability of the RL stack. The role requires proficiency in Python and C++, familiarity with Jax or Rust, and experience with large-scale RL/NLP training infrastructures and RL numerics. #J-18808-Ljbffr Xai

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the RL Infrastructure Engineer Scalable Training & Performance in Palo Alto, CA vacancy
  • xAI is seeking an engineer to join the RL infrastructure team to help develop our RL training framework. The RL engineer will design and implement systems...  ...debug and optimize end-to-end training performance and improve scalability and observability of the RL stack in a... 
    Training
    Performance

    Pantera Capital

    Palo Alto, CA
    1 day ago
  • $124k - $250k

     ...member of our software engineering infra team, you'll solve...  ...-of-the-art software infrastructure. The team builds a high-performance, high availability, globally...  ...and maintaining scalable infrastructure with high...  ...model delivery, including training, serving, and... 
    Training
    Performance

    AppLovin

    Palo Alto, CA
    1 day ago
  • $250k - $320k

     ...Staff Infrastructure Engineer We are partnered with a Stealth AI Lab (backed...  ...inference platform, GPU‑based training clusters, and data...  ...production—ensuring low‑latency performance, high availability, and...  ...pipelines for latency and scalability. A builder’s mindset — you... 
    Training
    Performance
    Full time
    Immediate start

    Strativ Group

    Menlo Park, CA
    2 days ago
  • $180k

     ...ML Infrastructure Engineer Palo Alto, California, United States SpaceXAI...  ...the reliable, high-performance ML platform that powers recommendations...  ...compute infrastructure, training frameworks, and...  ...across the stack Ensuring scalability, reliability, and efficiency... 
    Training
    Performance
    Temporary work
    Work experience placement

    Xai

    Palo Alto, CA
    3 days ago
  •  ...building the world’s most scalable driver, combining...  ...As a software engineering intern, you will work...  ...Onboard Systems, ML Infrastructure, Simulation, or Technical...  ...a reliable and high-performance platform that allows...  ...and optimize on-cloud training and onboard inference... 
    Training
    Performance
    Internship

    Nuro

    Mountain View, CA
    21 hours ago
  • $235.03k - $352.29k

     ...building the world’s most scalable driver, combining...  ..., the overall system performance depends heavily on...  ...and diversity of its training and evaluation data....  ...framework, supporting infrastructure, and a suite of data...  ...collaborates with autonomy engineers to ensure our labeled... 
    Training
    Performance
    Full time

    Nuro

    Mountain View, CA
    21 hours ago
  • $117.2k - $176.7k

     ...of the RoleThe Software Engineer (MTS) role is part of our...  ...team within the Cloud Infrastructure organization. Platform Engineering...  ...to the reliability, scalability, and developer...  ...compensation, promotion, benefits, training, assessment of job performance, discipline, termination... 
    Training
    Performance
    Full time

    Salesforce

    Palo Alto, CA
    4 days ago
  • $190k - $260k

     ...speed at which we can train it. Every improvement...  ...models - depends on infrastructure that turns thousands...  .... We are looking for engineers who make model training...  ...GPU computeDevelop scalable dataset construction...  ...offsExperience building high-performance data pipelines for... 
    Training
    Performance
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    1 day ago
  • $148.2k - $222.2k

    Staff Engineer, Infrastructure Platforms Position SummaryThe Staff Engineer,...  ...enterprise storage, High Performance Computing (HPC), and modern...  ...initiatives focused on automation, scalability, resiliency, and...  ...promoting collaboration, cross-training, and operational... 
    Training
    Performance
    Full time
    Work from home
    Monday to Friday

    Pacific Biosciences

    Menlo Park, CA
    21 hours ago
  • $235.03k - $352.29k

     ..., the overall system performance depends heavily on the...  ...and diversity of its training and evaluation data....  ...framework, supporting infrastructure, and a suite of data...  ...must be reliable and scalable. This includes everything...  ...with autonomy engineers to ensure our labeled... 
    Training
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  •  ...you to take your software engineering career to the next level....  ...to build and operate scalable, reliable ML training systems and pipelines on...  ...often GPU-based), improve performance and cost efficiency, and...  ...Build and operate training infrastructure on Kubernetes (e.g., EKS... 
    Training
    Performance

    JP Morgan Chase

    Palo Alto, CA
    3 days ago
  • $204k - $259k

     ...lifecycle, including pre-training and post-training....  ...We are looking for engineers with ML system expertise...  ...reinforcement learning (RL), building systems that...  ...for efficient and scalable learners/actors, and low...  ...quickly to improve model performance and training workflows... 
    Training
    Performance
    Full time
    Remote work

    Waymo

    Mountain View, CA
    21 hours ago
  • $170k - $235k

     ...on Mars.SR. SOFTWARE ENGINEER (PLATFORM TEAM) The Platform...  ...tooling and security infrastructure that empowers every...  ...team creates secure, scalable gateways and proxy systems...  ...to write code faster, perform advanced data analysis...  ...applications, and trained models on managed, reliable... 
    Training
    Performance
    Permanent employment
    Temporary work

    SpaceX

    Palo Alto, CA
    1 day ago
  • $158k - $237k

     ...The Forge Team: Engineering the Backbone of Rubrik...  ...reliable, secure, and scalable software-defined platform...  ...of the fundamental infrastructure, with a deep focus on...  ...clusters' health and performance. We build frameworks...  ...relevant education or training. US Pay Range $1... 
    Training
    Performance
    Full time
    Local area

    Rubrik

    Palo Alto, CA
    21 hours ago
  • $175k - $287k

     ...hybrid, meaning it will be performed both from home and from a...  ...LinkedIn’s AI model training, feature engineering and serving with hundreds...  ...user queries Model Training Infrastructure: As an engineer on the AI...  ...advance one of the most scalable AI platforms in the world... 
    Training
    Performance
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    3 days ago
  •  ...future of cloud platform engineering. As a Principal...  ...secure, reliable, and scalable. This is your opportunity...  ...expertise will drive performance, efficiency, and a...  ..., and reliable cloud infrastructure and platform tools.Drive...  ...clear documentation, training, office hours, and... 
    Training
    Performance
    Work at office
    Shift work

    JP Morgan Chase

    Palo Alto, CA
    1 day ago
  • $160.36k - $240.54k

     ...system, the overall system performance depends heavily on the...  ...and diversity of its training and evaluation data. The...  ...driving systems by creating a scalable and reliable data infrastructure. This infrastructure is...  ...closely with system engineers to thoroughly validate the... 
    Training
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  •  ...AI / ML Platform Engineers to build the platform...  ...workflows scalable, reliable, and reproducible...  ...focuses on the infrastructure and platform...  ...execution, distributed training and inference,...  ..., observability, performance, and developer...  ...research workflows for RL systems,... 
    Training
    Performance

    AMD

    Santa Clara, CA
    2 days ago
  • $160k - $225k

    Databricks is hiring a Senior Software Engineer for AI Runtime in Mountain...  ...of AIR's managed GPU training platform, ensuring scalable and resilient training. The ideal...  ...systems and strong knowledge of GPU performance and training infrastructure. This position comes with a... 
    Training
    Performance

    Databricks

    Mountain View, CA
    1 day ago
  • $180k - $260k

     ...researchers and veteran systems engineers who share a vision for...  ...complex, traditional infrastructure struggles to meet the demands of performance, reliability, and...  ..., and development of scalable network monitoring...  ...with NCCL, distributed training infrastructure, or AI... 
    Training
    Performance

    Clockwork.io

    Palo Alto, CA
    3 days ago
  • $150k - $217k

     ...technology velocity, systems performance, and cost reductions...  ...Google networks more scalable, efficient, and...  ...with other engineering teams.Lead and improve...  ...services.The AI and Infrastructure team is redefining what...  ...relevant education or training. US: $150000 - $21700... 
    Training
    Performance
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  • $88k - $121k

     ...'re searching for a Network Engineer I to join our Network Engineering...  ...and maintenance of network infrastructure (switches, routers, wireless...  ...network health and performance, responding to alerts, and helping...  ..., relevant education or training, and market conditions. These... 
    Training
    Performance
    Internship
    Work at office
    Local area
    3 days per week

    Aurora Innovation

    Mountain View, CA
    2 days ago
  • $217k - $312.2k

     ...Senior Engineering Manager for Workspace Platform Join...  ...world’s best data and AI infrastructure platform so our...  ...testing strategies, and performance optimizations for...  ...designing and building scalable distributed systems....  ...relevant certifications and training, and specific work... 
    Training
    Performance
    Local area
    Worldwide

    Databricks

    Mountain View, CA
    4 days ago
  • $274k - $304k

     ...AI Platform Engineer - Training & Inference Saviynt's AI-powered...  ...mission is to build a secure, scalable, product-agnostic AI...  ...sharing • Optimise inference performance: configure fractional GPU...  ...and cloud LLMs • Build RL training infrastructure: define Flyte workflows... 
    Training
    Performance

    Saviynt

    Milpitas, CA
    21 hours ago
  • $114.6k - $234.6k

     ...large-scale distributed infrastructure for the cloud? Oracle...  ...looking for hands-on engineers with expertise and...  ...architecting components of scalable, elastic distributed...  ...processing. -Design performance and load testing....  ...ongoing feedback and training to improve skills. Coaches... 
    Training
    Performance
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Santa Clara, CA
    21 hours ago
  •  ...Senior Lead Software Engineer Be an integral part...  ...Corporate Sector, Infrastructure Platforms team, you are...  ...secure, stable, and scalable way. Drive significant...  ...cloud resources for performance and cost efficiency....  ...skills Formal training or certification on software... 
    Training
    Performance
    For contractors

    Chase

    Palo Alto, CA
    21 hours ago
  • The RL infrastructure team is looking for an engineer to help develop our RL training framework Design and implement the systems backing all RL workloads at...  ...debug, and optimize end-to-end training performance Improve scalability and observability of the RL stack... 
    Training
    Performance
    Visa sponsorship
    Flexible hours

    Xai

    Palo Alto, CA
    3 days ago
  • $159.9k - $219.9k

     ...0 About Us: The Web Engineering team at Databricks builds...  ...on: framework, infrastructure, CI/CD, and deployment...  ...effort to a centralized, scalable system. The impact...  ...certifications and training, and specific work location...  ...for annual performance bonus, equity, and the... 
    Training
    Performance
    Worldwide
    Shift work

    Cacheflow

    Mountain View, CA
    1 day ago
  •  ...Network Development Engineer As a Principal Network...  ...within Oracle Cloud Infrastructure (OCI) Network...  ...ensure the reliability, scalability, and operational excellence...  ...the availability, performance, and reliability of OCI...  ...; mentoring, training, and supporting the development... 
    Training
    Performance

    Oracle

    Santa Clara, CA
    4 days ago
  • $153.2k - $234.1k

     ...As a Senior ML Infra Engineer, you will work on the...  ...dataset generation, training, evaluation and iteration...  ...pipelines that are performant, easy to use, and...  ...models that rely on the scalable, intuitive, and high‑...  ...applications, or ML infrastructure. Experience designing... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to RL Infrastructure Engineer Scalable Training & Performance. Be the first to apply!