RL Infrastructure Engineer Scalable Training & Performance
Xai
xAI in Palo Alto is seeking an experienced engineer to design and implement the RL training framework and the systems backing all RL workloads, from ablations to production runs. You will profile, debug, and optimize end-to-end training performance, and improve scalability and observability of the RL stack. The role requires proficiency in Python and C++, familiarity with Jax or Rust, and experience with large-scale RL/NLP training infrastructures and RL numerics. #J-18808-Ljbffr Xai
- xAI is seeking an engineer to join the RL infrastructure team to help develop our RL training framework. The RL engineer will design and implement systems... ...debug and optimize end-to-end training performance and improve scalability and observability of the RL stack in a...TrainingPerformance
$124k - $250k
...member of our software engineering infra team, you'll solve... ...-of-the-art software infrastructure. The team builds a high-performance, high availability, globally... ...and maintaining scalable infrastructure with high... ...model delivery, including training, serving, and...TrainingPerformance$250k - $320k
...Staff Infrastructure Engineer We are partnered with a Stealth AI Lab (backed... ...inference platform, GPU‑based training clusters, and data... ...production—ensuring low‑latency performance, high availability, and... ...pipelines for latency and scalability. A builder’s mindset — you...TrainingPerformanceFull timeImmediate start$180k
...ML Infrastructure Engineer Palo Alto, California, United States SpaceXAI... ...the reliable, high-performance ML platform that powers recommendations... ...compute infrastructure, training frameworks, and... ...across the stack Ensuring scalability, reliability, and efficiency...TrainingPerformanceTemporary workWork experience placement- ...building the world’s most scalable driver, combining... ...As a software engineering intern, you will work... ...Onboard Systems, ML Infrastructure, Simulation, or Technical... ...a reliable and high-performance platform that allows... ...and optimize on-cloud training and onboard inference...TrainingPerformanceInternship
$235.03k - $352.29k
...building the world’s most scalable driver, combining... ..., the overall system performance depends heavily on... ...and diversity of its training and evaluation data.... ...framework, supporting infrastructure, and a suite of data... ...collaborates with autonomy engineers to ensure our labeled...TrainingPerformanceFull time$117.2k - $176.7k
...of the RoleThe Software Engineer (MTS) role is part of our... ...team within the Cloud Infrastructure organization. Platform Engineering... ...to the reliability, scalability, and developer... ...compensation, promotion, benefits, training, assessment of job performance, discipline, termination...TrainingPerformanceFull time$190k - $260k
...speed at which we can train it. Every improvement... ...models - depends on infrastructure that turns thousands... .... We are looking for engineers who make model training... ...GPU computeDevelop scalable dataset construction... ...offsExperience building high-performance data pipelines for...TrainingPerformanceTemporary workWork at officeVisa sponsorship$148.2k - $222.2k
Staff Engineer, Infrastructure Platforms Position SummaryThe Staff Engineer,... ...enterprise storage, High Performance Computing (HPC), and modern... ...initiatives focused on automation, scalability, resiliency, and... ...promoting collaboration, cross-training, and operational...TrainingPerformanceFull timeWork from homeMonday to Friday$235.03k - $352.29k
..., the overall system performance depends heavily on the... ...and diversity of its training and evaluation data.... ...framework, supporting infrastructure, and a suite of data... ...must be reliable and scalable. This includes everything... ...with autonomy engineers to ensure our labeled...TrainingPerformanceImmediate startFlexible hours- ...you to take your software engineering career to the next level.... ...to build and operate scalable, reliable ML training systems and pipelines on... ...often GPU-based), improve performance and cost efficiency, and... ...Build and operate training infrastructure on Kubernetes (e.g., EKS...TrainingPerformance
$204k - $259k
...lifecycle, including pre-training and post-training.... ...We are looking for engineers with ML system expertise... ...reinforcement learning (RL), building systems that... ...for efficient and scalable learners/actors, and low... ...quickly to improve model performance and training workflows...TrainingPerformanceFull timeRemote work$170k - $235k
...on Mars.SR. SOFTWARE ENGINEER (PLATFORM TEAM) The Platform... ...tooling and security infrastructure that empowers every... ...team creates secure, scalable gateways and proxy systems... ...to write code faster, perform advanced data analysis... ...applications, and trained models on managed, reliable...TrainingPerformancePermanent employmentTemporary work$158k - $237k
...The Forge Team: Engineering the Backbone of Rubrik... ...reliable, secure, and scalable software-defined platform... ...of the fundamental infrastructure, with a deep focus on... ...clusters' health and performance. We build frameworks... ...relevant education or training. US Pay Range $1...TrainingPerformanceFull timeLocal area$175k - $287k
...hybrid, meaning it will be performed both from home and from a... ...LinkedIn’s AI model training, feature engineering and serving with hundreds... ...user queries Model Training Infrastructure: As an engineer on the AI... ...advance one of the most scalable AI platforms in the world...TrainingPerformanceFor contractorsWork experience placementWork at officeFlexible hours- ...future of cloud platform engineering. As a Principal... ...secure, reliable, and scalable. This is your opportunity... ...expertise will drive performance, efficiency, and a... ..., and reliable cloud infrastructure and platform tools.Drive... ...clear documentation, training, office hours, and...TrainingPerformanceWork at officeShift work
$160.36k - $240.54k
...system, the overall system performance depends heavily on the... ...and diversity of its training and evaluation data. The... ...driving systems by creating a scalable and reliable data infrastructure. This infrastructure is... ...closely with system engineers to thoroughly validate the...TrainingPerformanceImmediate startFlexible hours- ...AI / ML Platform Engineers to build the platform... ...workflows scalable, reliable, and reproducible... ...focuses on the infrastructure and platform... ...execution, distributed training and inference,... ..., observability, performance, and developer... ...research workflows for RL systems,...TrainingPerformance
$160k - $225k
Databricks is hiring a Senior Software Engineer for AI Runtime in Mountain... ...of AIR's managed GPU training platform, ensuring scalable and resilient training. The ideal... ...systems and strong knowledge of GPU performance and training infrastructure. This position comes with a...TrainingPerformance$180k - $260k
...researchers and veteran systems engineers who share a vision for... ...complex, traditional infrastructure struggles to meet the demands of performance, reliability, and... ..., and development of scalable network monitoring... ...with NCCL, distributed training infrastructure, or AI...TrainingPerformance$150k - $217k
...technology velocity, systems performance, and cost reductions... ...Google networks more scalable, efficient, and... ...with other engineering teams.Lead and improve... ...services.The AI and Infrastructure team is redefining what... ...relevant education or training. US: $150000 - $21700...TrainingPerformanceWorldwide$88k - $121k
...'re searching for a Network Engineer I to join our Network Engineering... ...and maintenance of network infrastructure (switches, routers, wireless... ...network health and performance, responding to alerts, and helping... ..., relevant education or training, and market conditions. These...TrainingPerformanceInternshipWork at officeLocal area3 days per week$217k - $312.2k
...Senior Engineering Manager for Workspace Platform Join... ...world’s best data and AI infrastructure platform so our... ...testing strategies, and performance optimizations for... ...designing and building scalable distributed systems.... ...relevant certifications and training, and specific work...TrainingPerformanceLocal areaWorldwide$274k - $304k
...AI Platform Engineer - Training & Inference Saviynt's AI-powered... ...mission is to build a secure, scalable, product-agnostic AI... ...sharing • Optimise inference performance: configure fractional GPU... ...and cloud LLMs • Build RL training infrastructure: define Flyte workflows...TrainingPerformance$114.6k - $234.6k
...large-scale distributed infrastructure for the cloud? Oracle... ...looking for hands-on engineers with expertise and... ...architecting components of scalable, elastic distributed... ...processing. -Design performance and load testing.... ...ongoing feedback and training to improve skills. Coaches...TrainingPerformanceTemporary workFlexible hoursShift work- ...Senior Lead Software Engineer Be an integral part... ...Corporate Sector, Infrastructure Platforms team, you are... ...secure, stable, and scalable way. Drive significant... ...cloud resources for performance and cost efficiency.... ...skills Formal training or certification on software...TrainingPerformanceFor contractors
- The RL infrastructure team is looking for an engineer to help develop our RL training framework Design and implement the systems backing all RL workloads at... ...debug, and optimize end-to-end training performance Improve scalability and observability of the RL stack...TrainingPerformanceVisa sponsorshipFlexible hours
$159.9k - $219.9k
...0 About Us: The Web Engineering team at Databricks builds... ...on: framework, infrastructure, CI/CD, and deployment... ...effort to a centralized, scalable system. The impact... ...certifications and training, and specific work location... ...for annual performance bonus, equity, and the...TrainingPerformanceWorldwideShift work- ...Network Development Engineer As a Principal Network... ...within Oracle Cloud Infrastructure (OCI) Network... ...ensure the reliability, scalability, and operational excellence... ...the availability, performance, and reliability of OCI... ...; mentoring, training, and supporting the development...TrainingPerformance
$153.2k - $234.1k
...As a Senior ML Infra Engineer, you will work on the... ...dataset generation, training, evaluation and iteration... ...pipelines that are performant, easy to use, and... ...models that rely on the scalable, intuitive, and high‑... ...applications, or ML infrastructure. Experience designing...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to RL Infrastructure Engineer Scalable Training & Performance. Be the first to apply!
- data infrastructure engineer Palo Alto, CA
- infrastructure engineering manager Palo Alto, CA
- senior infrastructure engineer Palo Alto, CA
- principal infrastructure engineer Palo Alto, CA
- remote infrastructure engineer Palo Alto, CA
- infrastructure developer Palo Alto, CA
- infrastructure engineer Palo Alto, CA
- performance testing Palo Alto, CA
- high performance computing engineer Palo Alto, CA
- performance test engineer Palo Alto, CA


