Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Agent Post-Training, Frontier Evals and Environments Research

United States Digital Space LLC

About the Team The Agent Post‑Training team creates the frontier agents the company ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi‑agent coordination, long‑horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what the company’s next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a researcher working on Frontier Evals & Environments, you will help build north star model environments to drive progress towards safe AGI/ASI. Your work will directly guide the research programs of the most ambitious training runs happening at the company. Some prior open‑source evaluations built by researchers in this role include GDPval, SWE‑bench Verified, MLE‑bench, PaperBench, and SWE‑Lancer. If you are interested in feeling firsthand the fast progress of our models, and steering them towards good outcomes, this is the role for you. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high‑agency role for people who want their work to land directly in frontier models. Responsibilities Create ambitious RL environments to push our models to their limits, and measure frontier model capabilities, skills, and behaviors. Develop new methodologies for automatically exploring the behavior of these models. Dive deep into the science of measurement, including understanding scalability, reliability, and variance of our evaluation methodology. Help steer training for our largest training runs, and see the future first. Design scalable systems and processes to support continuous evaluation. Build self‑improvement loops to automate model understanding. Qualifications Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across parts you have not worked in before. Have hands‑on experience with LLMs, RL, RLHF/RLAIF, post‑training, evals, graders, synthetic data, model training, coding agents, tool‑using agents, or production ML systems. Are excited by open‑ended problems where the path is unclear, the signal is noisy, and the right answer requires both research taste and engineering execution. Care about product impact and model behavior, not just benchmark movement; you have opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with. Can move from a vague behavioral problem to a concrete experiment: define the hypothesis, build the pipeline, run the model, analyze the result, and decide what to do next. Are comfortable working across research, product, infrastructure, data, evals, and safety boundaries, and can communicate clearly with each group. Like building load‑bearing systems and processes when that is what the team needs, even if the work is not glamorous. Want to train and ship the models that make agents genuinely useful for developers, enterprises, researchers, and everyday users. About the Company the company is an AI research and deployment company dedicated to ensuring that general‑purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. Equal Opportunity Employment We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see the company’s affirmative action and equal employment opportunity policy statement. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. Privacy Notice the company Global Applicant Privacy Policy #J-18808-Ljbffr United States Digital Space LLC

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Agent Post-Training, Frontier Evals and Environments Research in San Francisco, CA vacancy
  • OpenAI is seeking a Context Researcher on the Agent Post-Training team to scale compute spent on context and to...  ...model runs and ship enhancements to frontier products used by developers and...  ...You will design experiments, build evals and data pipelines, and contribute to... 
    Training

    Neura Market

    San Francisco, CA
    4 days ago
  • $380k

    Research Scientist - Multimodal Agent, Consumer Devices | OpenAI Careers Research Scientist...  .... We work at the frontier of multimodal AI, helping...  ...team to work on RLHF and post‑training for personalized, multimodal...  ...clean experiments, reliable evals, and decision‑useful... 
    Training
    Work at office
    Immediate start
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  • $180k - $225k

     ...transformation through frontier AI systems that...  ...production AI agents that automate...  ...architectures, and research papers emerge every...  ...in high-stakes environments.Collaborate with...  ...displayed on each job posting reflects the...  ...relevant education or training. Scale employees... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  • $58 - $63 per hour

    About the Role The Agents team investigates how to build, align, and scale frontier AI systems that can tackle complex, multi...  ...infrastructure—from designing post-training methods for agentic behavior to...  ...metrics fall short. As a research intern, you will work on problems... 
    Training
    Hourly pay
    Internship

    Togetherai

    San Francisco, CA
    15 hours ago
  • $190k - $270k

     ...TeamThe Databricks AI Research organization is pushing the frontier of next-generation...  ...the models and agents that unlock it. Our...  ...stack, from model training to advanced multi-agent...  ...pillars include post-training enhancements...  ...of specialized RL environments. As a member of this... 
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  •  ...building long-horizon multi-agent systems and pushing the boundaries of AI research, this is the kind of role...  ...novel problems at the frontier of what's currently possible...  ...-horizon reasoning, LLM post‑training for non‑myopic objectives, environment and feedback design.... 
    Training

    Techire Ai

    San Francisco, CA
    2 days ago
  • $218.4k - $273k

     ...accelerating the abundance of frontier data to pave the road...  ...the ACE team The Agent Capabilities & Environments (ACE) team, part of Scale’s Research organization, brings...  ...on each job posting reflects the minimum...  ...relevant education or training. Scale employees in eligible... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $216k - $270k

    Scale Labs, Research Scientist - Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral...  ...so by elements of their environment; Design and build exploits...  ...prototypes. Experience with post-training and RL techniques such as... 
    Training
    Full time

    Scale AI, Inc.

    San Francisco, CA
    1 day ago
  • Traverse is a research data lab building reinforcement learning environments for frontier AI labs. As a Research Scientist, you will design and build RL environments that...  ...and domain knowledge, turning that into training signals that actually work. We value rigorous... 
    Training

    Traverse

    San Francisco, CA
    3 days ago
  • $264.8k - $331k

     ...About the General Agents TeamThe General...  ...intersection of frontier agent development...  ...of real customer environments.You will:Design and...  ...spaces, balancing research-driven approaches...  ...on each job posting reflects the minimum...  ...relevant education or training. Scale employees... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $250k - $300k

     ...models at leading research labs and enterprises...  ...solutions for frontier AI development: Enterprise...  ...high-quality training data at scale...  ...Join Us High-Impact Environment : We operate like...  ...evaluate autonomous agent capabilities. Design...  ..., and blog posts. What You Bring Ph... 
    Training
    Work at office
    Flexible hours
    2 days per week

    Labelbox

    San Francisco, CA
    15 hours ago
  • Polymath Labs is hiring a Founding Member of Technical Staff - Research to push the frontier of autonomous agents. You will work on long-horizon evaluation, agent post-training, and environment design, wearing multiple hats from building benchmarks to running rigorous... 
    Training

    Polymath Labs

    San Francisco, CA
    1 day ago
  • $259.2k - $324k

     ...accelerating the abundance of frontier data to pave the road...  ...the ACE team The Agent Capabilities & Environments (ACE) team, part of Scale’s Research organization, brings...  ...on each job posting reflects the minimum...  ...relevant education or training. Scale employees in eligible... 
    Training
    Full time

    Scale AI, Inc.

    San Francisco, CA
    2 days ago
  • $252k - $315k

     ...systems in production environments.This is a Management role...  ...AI models and agents within customer security...  ...current on the latest AI/ML research and tools, bringing...  ...displayed on each job posting reflects the minimum and...  ...relevant education or training. Scale employees in... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...to-end software agents. We're the...  ...Role MissionPost-training is the critical...  ...role blends deep research and hands-on engineering...  ...across evals and production...  ...systems through post-training,...  ...moving research environments where priorities...  ...to operate at frontier scale from day... 
    Training
    Shift work

    Cognition AI

    San Francisco, CA
    3 days ago
  • United States Digital Space LLC is looking for a Context Researcher to scale compute on context in the Agent Post-Training team. This role involves collaborating with researchers and engineers on model training and improving product interfaces. Your responsibilities will... 
    Training

    United States Digital Space LLC

    San Francisco, CA
    2 days ago
  • Goaly is seeking an AI Researcher to lead research on agentic AI, focusing on training specialized models and building orchestration stacks. You'll design experiments, publish results, and advance the field. Ideal candidates have a Ph.D. or Master's in relevant fields,... 
    Training

    Goaly

    San Francisco, CA
    3 days ago
  • $264.8k - $331k

     ...from Meta, we are doubling down on building out state of the art post-training algorithms to reach the performance necessary for complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are working... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $25k

     ...agencies to safeguard the public through information sharing, training, research, and technology. Duties & Responsibilities This is an Open...  ...Announcement (OCA) to fill multiple Criminal Investigator (Special Agent) vacancies for the Homeland Security Task Force. Eligible... 
    Training
    Work at office
    Local area
    Immediate start
    Relocation package

    ATF

    San Francisco, CA
    4 days ago
  •  ...DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering...  ...overview As an AI Researcher, Agent, you will own the full...  ...stacks and simulated environments that make large‑scale...  ...from data curation through post‑training optimization.... 
    Training
    Full time
    Work at office

    Goaly

    San Francisco, CA
    3 days ago
  • $151.5k - $222.2k

     ...technologies and capabilities. Frontier AI is a purpose‑built...  ...models, multi‑agent systems, and robotics...  .... You will design the environments, rewards, and domain‑specific...  .... Responsibilities Research & Innovation: Partner...  ...scientific signal. Post‑train domain models (SFT, DPO... 
    Training
    Full time
    Flexible hours

    100 Eli Lilly and Company

    South San Francisco, CA
    2 days ago
  • $250k

    Research Scientist / Engineer - Multimodal Agent SF Bay Area, CA • Remote, International • London, UK | Research Remote • Hybrid Full-time About Luma AI Luma...  ...models, we believe that AI needs to be jointly trained over all signal modalities - text, video, audio, images... 
    Training
    Full time
    Remote work
    Worldwide

    Luma AI

    San Francisco, CA
    4 days ago
  • $216k - $270k

     ...Software Engineer, Agent Oversight About...  ...focused on pushing the frontier of what agentic...  ...paired with applied ML research, design, and...  ...tradeoffs between offline evals and live customer...  ...to own model training, modeling strategy...  ...on each job posting reflects the minimum... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $248.4k - $310.5k

     ...focused on pushing the frontier of what agentic...  ...paired with applied ML research, design, and evaluation...  ...Engineering Manager for Agent Oversight, you'll lead...  ...displayed on each job posting reflects the minimum and...  ...relevant education or training. Scale employees in eligible... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $205.6k - $257k

     ...advancements in AI, including frontier model training, enterprise adoption,...  ...vertical within our Agents Data & Reinforcement Learning Environments team. In this role,...  ...with a sense for AI research and current agent capabilities...  ...displayed on each job posting reflects the minimum... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...San Francisco is looking for a Member of Technical Staff - Research to advance autonomous agent capabilities. The ideal candidate will engage in core research areas, including agent evaluation and environment design, while also contributing to building benchmarks and writing... 

    Polymath

    San Francisco, CA
    5 days ago
  •  ...the Team API Agents builds the shared...  ...that turn OpenAI’s frontier models into...  ...software engineering, research, finance, healthcare...  ..., execution environments, identity and permissions...  ...than model training: success comes from...  ...believe this job posting is non-compliant,... 
    Training
    Full time
    Internship

    OpenAI

    San Francisco, CA
    1 day ago
  • Scale is seeking a Machine Learning Research Engineer, Agents - Enterprise GenAI to advance state-of-the-art Agent RL training for enterprise datasets. You will train models, iterate...  ...with LLMs in production, experience with post-training methods (RLHF/RLVR, PPO/GRPO), and... 
    Training

    Scale

    San Francisco, CA
    3 days ago
  •  ...About the Team OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Agent Harness Bridge team sits...  ...effectively in training environments. This role is ideal...  ...you believe this job posting is non-compliant, please... 
    Training
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $110k - $120k

     ...built on a diverse culture of trained, talented, and mission-...  ...protective services program with a frontier technology client with...  ...full-time Executive Protection Agents with high integrity, sound judgment...  ...maintain a safe and secure environment. Operate in a low-profile, low... 
    Training
    Hourly pay
    Daily paid
    Full time
    Temporary work
    Local area
    Relocation package
    Flexible hours
    Shift work
    Night shift
    Weekend work

    Surefox North America Inc

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Agent Post-Training, Frontier Evals and Environments Research. Be the first to apply!