Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Frontier AI Evaluation Engineer: RL Environments

OpenAI

A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models and AGI/ASI measurement. Candidates should have strong engineering and statistical skills, a red-teaming mindset, and the ability to thrive in a fast-paced research environment. This role offers the opportunity to shape innovative AI measurements and build evaluation environments that drive progress. #J-18808-Ljbffr OpenAI

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Frontier AI Evaluation Engineer: RL Environments in San Francisco, CA vacancy
  • HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning environments with verifiable rewards for real-world use cases, creating...  ..., reward functions, and evaluation harnesses to measure model performance... 
    Suggested

    HeyMilo AI

    San Francisco, CA
    1 day ago
  • Jack & Jill in San Francisco builds RL environments and robust backend services to stress-test frontier AI models. You will collaborate with Applied Researchers to translate...  ...designs into production-grade software and evaluation systems. The role demands hands-on work with... 
    Suggested

    Jack & Jill

    San Francisco, CA
    4 days ago
  •  ...is hiring a junior software engineer to design and build reinforcement-learning environments for software engineering tasks...  ...full lifecycle from concept to evaluation against frontier models, using Python and...  ...applying strong fundamentals to AI model behavior, and you will... 
    Suggested

    Simplify

    San Francisco, CA
    1 day ago
  • $200k

     ...builds the training data and evaluation infrastructure that frontier AI labs use to make their...  ...The Role As a SWE (Environments), you will design the datasets...  ...for/interned for any RL environment companies in...  ...Former founders and early engineers at early stage startups are... 
    Suggested
    Full time

    AfterQuery

    San Francisco, CA
    23 hours ago
  • CoffeeSpace is recruiting a Platform Engineer in San Francisco to build the infrastructure and tooling for frontier RL environments in financial services. You will own significant parts of the ML infrastructure, collaborate with ML engineers and product teams, and help... 
    Suggested
    Work at office
    Visa sponsorship

    CoffeeSpace

    San Francisco, CA
    3 days ago
  • $252k - $315k

    About Scale AI At Scale, our mission is to develop...  ...real impact. Scale Frontier Data is the organization...  ...the training and evaluation data that frontier labs...  ...Reinforcement learning environments are now the center of...  ...Responsibilities As a Staff Software Engineer, RL Environments, you'll... 
    Full time
    Work experience placement
    Remote work

    Scale AI

    San Francisco, CA
    2 days ago
  • $245k - $300k

     ...virtuous cycle: human insight improves AI, and better AI expands what people can...  ...into the data , evals , and RL environments frontier models learn from. We work with leading...  ...it's running. You sit between Pareto's engineering team and the researchers at the labs we... 
    Full time

    Pareto B.v.

    San Francisco, CA
    23 hours ago
  • Roboflow is hiring a Member of Technical Staff for the Frontier Data team. You will build RL environments, evaluations, and datasets that close high-value capability...  .... You’ll work directly with researchers and engineers to shape dataset development and model roadmaps.... 

    Roboflow

    San Francisco, CA
    13 hours ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI technology...  ...operations for the frontier of AI. Our customers include...  ...AI systems using Evaluation-Driven Development—an...  ...production.AI Evaluation Engineers focus on designing and...  ...run in production environments. You treat evaluation... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    1 day ago
  •  ...DesignArena, invites you to join a talent-dense team in San Francisco with 5.5M+ users and a rapidly growing platform. You will define how frontier AI models are measured, design new benchmarks, run experiments, and publish analyses that become industry gold standards. We sponsor... 
    Relocation
    Visa sponsorship

    Intelligence

    San Francisco, CA
    4 days ago
  • OpenAI is seeking an exceptional health AI researcher to build frontier capabilities that translate into real-world...  ...candidate will contribute to pretraining, RL/post-training, and evaluation, working with researchers, engineers, clinicians, and product teams to deliver... 

    AI Chopping Block

    San Francisco, CA
    13 hours ago
  • RippleMatch Inc. is seeking an innovative and motivated individual to design and refine reinforcement learning tasks in San Francisco. This role requires a strong command of Python and the ability to work independently with coding agents. Responsibilities include the full...

    RippleMatch

    San Francisco, CA
    2 days ago
  • $130k - $220k

     ...the leading independent AI benchmarking and...  ...insights company. They help engineers, enterprises,...  ...benchmarks do not just measure frontier AI — they actively...  ...best described as an AI Evaluation Engineer / Technical Generalist...  ...a fast-moving startup environment. **Tech Stack** *... 
    Full time
    Worldwide

    Aurora Jobs ApS

    San Francisco, CA
    a month ago
  • Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems... 

    Obsidian

    San Francisco, CA
    3 days ago
  •  ...Remote Industry: Applied AI / AI research data...  ...quality reinforcement-learning environments and agents sold to the world...  ...The Opportunity As an RL Environment Software Engineer, you will sit at the intersection...  ...-quality environments the frontier will need next, then build... 
    Remote work

    talentpluto

    San Francisco, CA
    7 days ago
  • Fundamental AI Research Institute in San Francisco is seeking researchers experienced in LLM post-training, agentic infrastructure, RL environments, and evaluation harnesses to accelerate flagship AI research. Responsibilities include building and scaling post-training... 

    Storm3

    San Francisco, CA
    3 days ago
  •  ...The Post-Training Frontiers team creates the frontier...  ...go into the final RL run and deciding...  ...will work across engineering and infrastructure...  ...data, RL systems, evaluation infrastructure,...  ...by fast-moving environments where reliability,...  ...OpenAI OpenAI is an AI research and... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...innovative tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and technical... 

    Scale AI

    San Francisco, CA
    1 day ago
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes,... 

    Mercor Inc

    San Francisco, CA
    13 hours ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering...  ...You will author original, executable research problems that frontier models cannot solve. Engagement lasts 6 weeks, part-time with... 
    Part time
    Immediate start

    Mercor

    San Francisco, CA
    13 hours ago
  •  ...Scientist to design novel benchmarks and evaluate frontier language models and agents. You will...  ...experiments and collaborate with research engineers, foundation-model developers and domain...  ...and present findings that influence how leading AI #J-18808-Ljbffr Cerebro
    Relocation

    Cerebro

    San Francisco, CA
    2 days ago
  • Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations... 

    Scale

    San Francisco, CA
    4 days ago
  •  ...seeking an exceptional researcher to build frontier health capabilities and translate them...  ...products. The candidate should have strong ML/AI research depth, be a hands‑on builder, and...  ...a track record in advancing pretraining, RL/post‑training, or health‑focused AI problems... 

    OpenAI

    San Francisco, CA
    4 days ago
  • $200k - $275k

    Halluminate is seeking a Platform Engineer to own our platform engineering, develop frontier RL environments, and help scale our engineering team from the ground up. The role is based in San Francisco with 5 days in the office, offering a base salary of $200-275k plus... 
    Work at office
    Relocation package

    Halluminate

    San Francisco, CA
    2 days ago
  • Obsidian is seeking experienced AI Safety Practitioners to evaluate frontier AI models for safety, quality, and alignment on complex topics. You will assess responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. Responsibilities... 

    Obsidian

    San Francisco, CA
    13 hours ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against... 
    Part time
    Immediate start

    Obsidian

    San Francisco, CA
    2 days ago
  • $240k - $280k

    A leading software monitoring company is seeking a Senior Software Engineer on its AI/ML team to build evaluation infrastructure for measuring the performance of AI systems. This role involves designing datasets, creating benchmarks, and ensuring AI features behave reliably... 

    Sentry

    San Francisco, CA
    23 hours ago
  •  ...Description Accellor is an AI-native services firm...  ...advanced AI, data, and engineering capabilities. Our...  ...engineering, cost optimization, evaluation gates, observability,...  ...velocity. Support frontier model workflows across...  ...distributed training, RL infrastructure, checkpointing... 

    Accellor

    San Francisco, CA
    14 days ago
  •  ...intelligence to power the AI economy. We partner with leading...  ...talent network trains frontier AI models in the same way...  ...As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at...  ...synthetic data pipelines and environments that generate high-signal... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    23 hours ago
  • Perplexity AI, Inc. is seeking energetic engineers to join our Agents engineering team. You will work across backend, full-stack, and AI/ML to build harnesses...  ...to solve open problems in AI and advance the frontier of agent capabilities for millions of users. A strong eye... 

    Perplexity AI, Inc.

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Frontier AI Evaluation Engineer: RL Environments. Be the first to apply!