Research Scientist - Model-Based Reinforcement Learning
Phaidra
Who You Are
Research Scientists at Phaidra lead our efforts in developing novel algorithmic architecture towards the end goal of bringing intelligent control systems to the industrial sector.
Having pioneered research in the world's leading academic and industrial labs as PhDs, post-docs or professors, Research Scientists join Phaidra to work collaboratively within and across research fields.
Drawing on expertise from a variety of disciplines including model-based reinforcement learning , planning and optimal control, deep learning, world models, and safe reinforcement learning, our Research Scientists are at the forefront of groundbreaking research and applying it to real-world industrial control systems.
Responsibilities
- Design, implement, and evaluate model-based reinforcement learning agents — including planning-based controllers ( MPC, MPPI ) — and the software prototypes needed to deploy them on real industrial control systems.
- Develop learned dynamics and world models (learned surrogates) that generalize across systems, including the training pipelines — pretraining, curriculum learning, active/adversarial learning, and fine-tuning — needed to make them reliable for planning and control.
- Research and implement methods for e.g. safe RL , constrained control, scenario planning and Bayesian RL , to develop agents that satisfy safety constraints during deployment.
- Report and present research findings and developments including status and results clearly and efficiently both internally and externally, verbally and in writing.
- Participate in and organize ambitious collaborative research projects, and work with external collaborators and partners to translate research into production outcomes.
- Mentor and guide Research Engineers to apply research findings and developments to industrial domains.
- Independently defines new research directions
- Translates research into practical outcomes
- Owns the development and rollout for an entire research area or large project
Key Qualifications
- PhD in a technical field or equivalent practical experience, with a strong background in model-based reinforcement learning and demonstrated knowledge in one or more of the following:
- Planning algorithms
- World models / learned dynamics surrogates
- Reinforcement Learning and Deep Learning
- Control Theory
- Safe / constrained RL
- Either:
- 2+ years of research experience in academia or industry after PhD graduation.
- 5+ years of research experience in academia or industry after Master’s graduation.
- Extensive research in the fields of {ModelBased, ModelFree, Safe} RL and Control Theory, with particular depth in model-based methods.
- Hands-on experience building and evaluating agents against simulators (e.g. differentiable simulators or world models) and closing the sim-to-real gap.
- Alignment with Phaidra's values: Agency, Velocity, Craft, & Truth .
Preferred Skills & Experience
- PhD in machine learning, control, or a closely related field.
- Deep, hands-on experience with model-based RL and planning agents applied to real-world dynamical or industrial systems.
- Strong Python and PyTorch skills, including vectorized/differentiable simulators and scaling experiments on distributed compute (e.g. Ray, Kubernetes, GCP ).
- A proven track record of publications in RL, control, or a related area.
- A real passion for AI and applying it to industrial systems to improve resource efficiency.
Our Stack
- Python
- PyTorch, scipy
- Kubernetes, Docker, Ray
- GCP
Onboarding
In your first 30 days...
- You will be immersed in an onboarding program that introduces you to Phaidra and our product.
- You will spend time in the research team and get introduced to the different tracks of research that we perform.
- You will learn how other teams operate, interact, and approach problems.
- You will read various parts of our handbook and familiarize yourself with the documentation culture at Phaidra.
- You will set up your development environment and get introduced to various parts of our code base.
By your first 60 days...
- You will have a solid understanding of what Phaidra does and how we do it.
- You will have met with team members across Phaidra and started building relationships that will help you be successful at your job.
- You will have a good understanding of our research tracks, which allowed you to form a plan for your first project, or maybe, you have already started one.
By your first 90 days...
- You will have been fully integrated in the team and with team members across the company.
- You will have started to contribute to knowledge sharing throughout Phaidra.
- Your project is in full swing and you might have some early results already.
£140 per hour
...technical deliverables that help train and evaluate advanced AI models. This short, intensive project brings together specialists across... ...£140 to £200 per hour. Eligibility ~ You must be based in the United Kingdom and have the right to work in the United Kingdom...SuggestedHourly payRemote work- ...Our advanced AI software and foundation models enable vehicles to perceive, understand,... ...in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter... .... Experience operating cloud-based services using Kubernetes, with a good understanding...SuggestedFull timeWork at officeWork from home
- ...The role We’re looking for a Model Release Engineer to build the platform that moves... ...priorities. Experience operating cloud-based services using Kubernetes, with a good understanding... ...Experience with MLOps, machine-learning infrastructure or production model-...SuggestedFull timeWork at officeWork from home
$50 - $70 per hour
...documents, slides, and spreadsheets. Previous AI or machine learning experience is not required. Guidelines and context will be provided... ...Work Terms Remote, hourly engagement. Applicants must be based in the United Kingdom or Europe. Compensation ~ Hourly...SuggestedHourly payRemote work£140 per hour
...documents and deliverables that help train and evaluate advanced AI models. You will collaborate with specialists from two other domains,... ...£140 to £200 per hour. Eligibility You must be based in the United Kingdom. You must have the right to work in the...SuggestedHourly payRemote work- ...We are seeking a Senior Scientist to design peptide ligands for delivery of RNA therapeutics... ...RFpeptide and BoltzGen. Perform physics-based filtering of designs via molecular... ...effectively to leadership. Enthusiasm to learn the fundamentals of RNA therapeutics....Live in
$800 per month
...differentiation, incorrect policy-term usage, and unsuitable acceptance or declination decisions. Provide written feedback to improve model behavior. Participate in onboarding office hours and calibration sessions. Qualifications At least 2 years of professional...For contractorsWork at officeImmediate startRemote work£50 - £100 per hour
...natural, expressive text-to-speech voices for an internal customer-experience AI agent. This remote recording opportunity is for UK-based native English speakers with a neutral, standard UK English accent who can deliver clear, engaging speech across a range of scripts....Hourly payRemote work£141k - £148k per year
...case Travel Requirements ~ Travel is required for key SDO committee meetings, conferences, and team onsites. The expected base salary range for this full-time position is listed below. Actual starting pay will be based on job-related factors, including exact...Full time$80 per hour
...product-type or tax-treatment errors, and missed replacement or disclosure requirements. Provide written feedback used to improve model behavior. Participate in onboarding office hours and calibration sessions. Qualifications At least 2 years of professional...Work at officeImmediate startRemote work- ...help develop foundational generative AI models. You will strengthen the quality of AI training... ...Key Responsibilities Partner with research and engineering teams to identify gaps in... ..., Golden Gate, Gateway, and restriction-based cloning, as well as codon optimization....Hourly payRemote workWeekday work
$50 - $100 per hour
...evaluate a next generation text to speech system for an internal customer experience AI agent. This opportunity is for female speakers based in the northern UK with authentic regional accents such as Yorkshire, Manchester, Newcastle, or Liverpool. Recordings must be...Hourly payRemote work10 hours per week$50 - $100 per hour
...emotion, and style when requested. Qualifications Native English speaker with a neutral, standard UK English accent, currently based in the United Kingdom. Proven experience in voice acting, dubbing, narration, podcasting, or broadcasting. Access to a...Hourly payRemote work10 hours per week- ...AI Security Researcher / Research Engineer Remote or hybrid - London preferred but North America, EU, UK accepted About Brave Brave... ...g., prompt injection, tool abuse, data exfiltration) Analyze model behavior (including internal thoughts) to understand and forecast...Full timeWork experience placementRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - Model-Based Reinforcement Learning. Be the first to apply!
- remote legal research United Kingdom
- research and development analyst United Kingdom
- research and development manager United Kingdom
- research and development chef United Kingdom
- pharmacovigilance scientist
- regulatory scientist
- applied scientist nlp
- quantum computing scientist
- entry level pharmaceutical scientist
- pain research scientist


