Frontier AI Evaluation Engineer: RL Environments
OpenAI
A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models and AGI/ASI measurement. Candidates should have strong engineering and statistical skills, a red-teaming mindset, and the ability to thrive in a fast-paced research environment. This role offers the opportunity to shape innovative AI measurements and build evaluation environments that drive progress. #J-18808-Ljbffr OpenAI
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Frontier AI Evaluation Engineer: RL Environments in San Francisco, CA vacancy
- HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning environments with verifiable rewards for real-world use cases, creating... ..., reward functions, and evaluation harnesses to measure model performance...Suggested
- Jack & Jill in San Francisco builds RL environments and robust backend services to stress-test frontier AI models. You will collaborate with Applied Researchers to translate... ...designs into production-grade software and evaluation systems. The role demands hands-on work with...Suggested
- ...is hiring a junior software engineer to design and build reinforcement-learning environments for software engineering tasks... ...full lifecycle from concept to evaluation against frontier models, using Python and... ...applying strong fundamentals to AI model behavior, and you will...Suggested
$200k
...builds the training data and evaluation infrastructure that frontier AI labs use to make their... ...The Role As a SWE (Environments), you will design the datasets... ...for/interned for any RL environment companies in... ...Former founders and early engineers at early stage startups are...SuggestedFull time- CoffeeSpace is recruiting a Platform Engineer in San Francisco to build the infrastructure and tooling for frontier RL environments in financial services. You will own significant parts of the ML infrastructure, collaborate with ML engineers and product teams, and help...SuggestedWork at officeVisa sponsorship
$252k - $315k
About Scale AI At Scale, our mission is to develop... ...real impact. Scale Frontier Data is the organization... ...the training and evaluation data that frontier labs... ...Reinforcement learning environments are now the center of... ...Responsibilities As a Staff Software Engineer, RL Environments, you'll...Full timeWork experience placementRemote work$245k - $300k
...virtuous cycle: human insight improves AI, and better AI expands what people can... ...into the data , evals , and RL environments frontier models learn from. We work with leading... ...it's running. You sit between Pareto's engineering team and the researchers at the labs we...Full time- Roboflow is hiring a Member of Technical Staff for the Frontier Data team. You will build RL environments, evaluations, and datasets that close high-value capability... .... You’ll work directly with researchers and engineers to shape dataset development and model roadmaps....
$150k - $250k
About Distyl AI Distyl is an applied AI technology... ...operations for the frontier of AI. Our customers include... ...AI systems using Evaluation-Driven Development—an... ...production.AI Evaluation Engineers focus on designing and... ...run in production environments. You treat evaluation...Work at office3 days per week- ...DesignArena, invites you to join a talent-dense team in San Francisco with 5.5M+ users and a rapidly growing platform. You will define how frontier AI models are measured, design new benchmarks, run experiments, and publish analyses that become industry gold standards. We sponsor...RelocationVisa sponsorship
- OpenAI is seeking an exceptional health AI researcher to build frontier capabilities that translate into real-world... ...candidate will contribute to pretraining, RL/post-training, and evaluation, working with researchers, engineers, clinicians, and product teams to deliver...
- RippleMatch Inc. is seeking an innovative and motivated individual to design and refine reinforcement learning tasks in San Francisco. This role requires a strong command of Python and the ability to work independently with coding agents. Responsibilities include the full...
$130k - $220k
...the leading independent AI benchmarking and... ...insights company. They help engineers, enterprises,... ...benchmarks do not just measure frontier AI — they actively... ...best described as an AI Evaluation Engineer / Technical Generalist... ...a fast-moving startup environment. **Tech Stack** *...Full timeWorldwide- Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems...
- ...Remote Industry: Applied AI / AI research data... ...quality reinforcement-learning environments and agents sold to the world... ...The Opportunity As an RL Environment Software Engineer, you will sit at the intersection... ...-quality environments the frontier will need next, then build...Remote work
- Fundamental AI Research Institute in San Francisco is seeking researchers experienced in LLM post-training, agentic infrastructure, RL environments, and evaluation harnesses to accelerate flagship AI research. Responsibilities include building and scaling post-training...
- ...The Post-Training Frontiers team creates the frontier... ...go into the final RL run and deciding... ...will work across engineering and infrastructure... ...data, RL systems, evaluation infrastructure,... ...by fast-moving environments where reliability,... ...OpenAI OpenAI is an AI research and...Full time
- ...innovative tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and technical...
- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes,...
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering... ...You will author original, executable research problems that frontier models cannot solve. Engagement lasts 6 weeks, part-time with...Part timeImmediate start
- ...Scientist to design novel benchmarks and evaluate frontier language models and agents. You will... ...experiments and collaborate with research engineers, foundation-model developers and domain... ...and present findings that influence how leading AI #J-18808-Ljbffr CerebroRelocation
- Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations...
- ...seeking an exceptional researcher to build frontier health capabilities and translate them... ...products. The candidate should have strong ML/AI research depth, be a hands‑on builder, and... ...a track record in advancing pretraining, RL/post‑training, or health‑focused AI problems...
$200k - $275k
Halluminate is seeking a Platform Engineer to own our platform engineering, develop frontier RL environments, and help scale our engineering team from the ground up. The role is based in San Francisco with 5 days in the office, offering a base salary of $200-275k plus...Work at officeRelocation package- Obsidian is seeking experienced AI Safety Practitioners to evaluate frontier AI models for safety, quality, and alignment on complex topics. You will assess responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. Responsibilities...
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against...Part timeImmediate start
$240k - $280k
A leading software monitoring company is seeking a Senior Software Engineer on its AI/ML team to build evaluation infrastructure for measuring the performance of AI systems. This role involves designing datasets, creating benchmarks, and ensuring AI features behave reliably...- ...Description Accellor is an AI-native services firm... ...advanced AI, data, and engineering capabilities. Our... ...engineering, cost optimization, evaluation gates, observability,... ...velocity. Support frontier model workflows across... ...distributed training, RL infrastructure, checkpointing...
- ...intelligence to power the AI economy. We partner with leading... ...talent network trains frontier AI models in the same way... ...As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at... ...synthetic data pipelines and environments that generate high-signal...Full timeWork at officeRelocation package
- Perplexity AI, Inc. is seeking energetic engineers to join our Agents engineering team. You will work across backend, full-stack, and AI/ML to build harnesses... ...to solve open problems in AI and advance the frontier of agent capabilities for millions of users. A strong eye...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Frontier AI Evaluation Engineer: RL Environments. Be the first to apply!
Related searches
- ai ml engineer San Francisco, CA
- ai prompt engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai developer San Francisco, CA
- ai engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- environment artist San Francisco, CA
- project manager environment San Francisco, CA


