Research Scientist - Model-Based Reinforcement Learning
Phaidra
Who You Are
Research Scientists at Phaidra lead our efforts in developing novel algorithmic architecture towards the end goal of bringing intelligent control systems to the industrial sector.
Having pioneered research in the world's leading academic and industrial labs as PhDs, post-docs or professors, Research Scientists join Phaidra to work collaboratively within and across research fields.
Drawing on expertise from a variety of disciplines including model-based reinforcement learning , planning and optimal control, deep learning, world models, and safe reinforcement learning, our Research Scientists are at the forefront of groundbreaking research and applying it to real-world industrial control systems.
Responsibilities
- Design, implement, and evaluate model-based reinforcement learning agents — including planning-based controllers ( MPC, MPPI ) — and the software prototypes needed to deploy them on real industrial control systems.
- Develop learned dynamics and world models (learned surrogates) that generalize across systems, including the training pipelines — pretraining, curriculum learning, active/adversarial learning, and fine-tuning — needed to make them reliable for planning and control.
- Research and implement methods for e.g. safe RL , constrained control, scenario planning and Bayesian RL , to develop agents that satisfy safety constraints during deployment.
- Report and present research findings and developments including status and results clearly and efficiently both internally and externally, verbally and in writing.
- Participate in and organize ambitious collaborative research projects, and work with external collaborators and partners to translate research into production outcomes.
- Mentor and guide Research Engineers to apply research findings and developments to industrial domains.
- Independently defines new research directions
- Translates research into practical outcomes
- Owns the development and rollout for an entire research area or large project
Key Qualifications
- PhD in a technical field or equivalent practical experience, with a strong background in model-based reinforcement learning and demonstrated knowledge in one or more of the following:
- Planning algorithms
- World models / learned dynamics surrogates
- Reinforcement Learning and Deep Learning
- Control Theory
- Safe / constrained RL
- Either:
- 2+ years of research experience in academia or industry after PhD graduation.
- 5+ years of research experience in academia or industry after Master’s graduation.
- Extensive research in the fields of {ModelBased, ModelFree, Safe} RL and Control Theory, with particular depth in model-based methods.
- Hands-on experience building and evaluating agents against simulators (e.g. differentiable simulators or world models) and closing the sim-to-real gap.
- Alignment with Phaidra's values: Agency, Velocity, Craft, & Truth .
Preferred Skills & Experience
- PhD in machine learning, control, or a closely related field.
- Deep, hands-on experience with model-based RL and planning agents applied to real-world dynamical or industrial systems.
- Strong Python and PyTorch skills, including vectorized/differentiable simulators and scaling experiments on distributed compute (e.g. Ray, Kubernetes, GCP ).
- A proven track record of publications in RL, control, or a related area.
- A real passion for AI and applying it to industrial systems to improve resource efficiency.
Our Stack
- Python
- PyTorch, scipy
- Kubernetes, Docker, Ray
- GCP
Onboarding
In your first 30 days...
- You will be immersed in an onboarding program that introduces you to Phaidra and our product.
- You will spend time in the research team and get introduced to the different tracks of research that we perform.
- You will learn how other teams operate, interact, and approach problems.
- You will read various parts of our handbook and familiarize yourself with the documentation culture at Phaidra.
- You will set up your development environment and get introduced to various parts of our code base.
By your first 60 days...
- You will have a solid understanding of what Phaidra does and how we do it.
- You will have met with team members across Phaidra and started building relationships that will help you be successful at your job.
- You will have a good understanding of our research tracks, which allowed you to form a plan for your first project, or maybe, you have already started one.
By your first 90 days...
- You will have been fully integrated in the team and with team members across the company.
- You will have started to contribute to knowledge sharing throughout Phaidra.
- Your project is in full swing and you might have some early results already.
- ...software and foundation models enable vehicles... ..., constantly learning and evolving as... ...looking for Applied Scientists to join Wayve... ...High-conviction Research Team With The Strategic... ...(e.g., diffusion-based, autoregressive,... ...Advance Reinforcement Learning and Reward...SuggestedFull timeWork at officeWork from homeVisa sponsorshipRelocation packageFlexible hours
$70 - $105 per hour
...training data for advanced generative AI models. You will design and review DNA and RNA constructs... ...solutions and rubrics, and work with research and engineering teams to close biological... ...employer and does not discriminate based on legally protected characteristics....SuggestedRemote jobHourly payFull timePart timeWeekday work$50 - $70 per hour
...documents, slides, and spreadsheets. Previous AI or machine learning experience is not required. Guidelines and context will be provided... ...Work Terms Remote, hourly engagement. Applicants must be based in the United Kingdom or Europe. Compensation ~ Hourly...SuggestedHourly payRemote work- ...Are Odyssey is an AI lab pioneering general world models: causal, multimodal systems that learn to predict and interact with the world over long horizons... ...cars. They’ve now brought together a world-class research team from DeepMind, Tesla, Waymo, Meta, Apple, and Wayve...SuggestedFull time
$150 per hour
...Overview Help evaluate how advanced language models reason through complex psychiatric cases.... .... Collaborate directly with AI research teams. Provide feedback on recurring failure... ...or open disciplinary action. Must be based in the United States, United Kingdom,...SuggestedHourly payPrivate practiceRemote work10 hours per week- ...clinical tasks to help advance frontier AI models across a range of clinical scenarios and... ...Licensure: Registered doctor based in the US, Europe or Asia-Pacific. Experience... ...are evaluated for clinical safety, working alongside clinicians and AI researchers....Hourly payContract workTemporary workPart timeFor contractorsRemote workFlexible hours
$50 - $100 per hour
...multiple takes with differences in emphasis, emotion, and style when requested Qualifications Native English speaker, currently based in the United Kingdom, with a neutral, standard UK English accent Proven experience in voice acting, dubbing, narration,...Hourly payRemote work10 hours per week$150 per hour
...on the evaluation of how large language models process complex psychiatric cases. You will... ...authoritative, gold-standard answers based on your expertise. Key Responsibilities... ...evaluations. Collaborate directly with research teams to enhance model performance. Provide...Private practice$50 - $100 per hour
...evaluate a next generation text to speech system for an internal customer experience AI agent. This opportunity is for female speakers based in the northern UK with authentic regional accents such as Yorkshire, Manchester, Newcastle, or Liverpool. Recordings must be...Hourly payRemote work10 hours per week$50 - $100 per hour
...emotion, and style when requested. Qualifications Native English speaker with a neutral, standard UK English accent, currently based in the United Kingdom. Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting. Access to a...Hourly payRemote work10 hours per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - Model-Based Reinforcement Learning. Be the first to apply!
- graduate research intern United Kingdom
- research and development internship United Kingdom
- research and development analyst United Kingdom
- remote legal research United Kingdom
- research and development chef United Kingdom
- research and development manager United Kingdom
- research writer United Kingdom
- scientist ii
- upstream scientist
- drug discovery scientist



