Agent Post-Training, Frontier Evals and Environments Research
United States Digital Space LLC
About the Team The Agent Post‑Training team creates the frontier agents the company ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi‑agent coordination, long‑horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what the company’s next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a researcher working on Frontier Evals & Environments, you will help build north star model environments to drive progress towards safe AGI/ASI. Your work will directly guide the research programs of the most ambitious training runs happening at the company. Some prior open‑source evaluations built by researchers in this role include GDPval, SWE‑bench Verified, MLE‑bench, PaperBench, and SWE‑Lancer. If you are interested in feeling firsthand the fast progress of our models, and steering them towards good outcomes, this is the role for you. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high‑agency role for people who want their work to land directly in frontier models. Responsibilities Create ambitious RL environments to push our models to their limits, and measure frontier model capabilities, skills, and behaviors. Develop new methodologies for automatically exploring the behavior of these models. Dive deep into the science of measurement, including understanding scalability, reliability, and variance of our evaluation methodology. Help steer training for our largest training runs, and see the future first. Design scalable systems and processes to support continuous evaluation. Build self‑improvement loops to automate model understanding. Qualifications Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across parts you have not worked in before. Have hands‑on experience with LLMs, RL, RLHF/RLAIF, post‑training, evals, graders, synthetic data, model training, coding agents, tool‑using agents, or production ML systems. Are excited by open‑ended problems where the path is unclear, the signal is noisy, and the right answer requires both research taste and engineering execution. Care about product impact and model behavior, not just benchmark movement; you have opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with. Can move from a vague behavioral problem to a concrete experiment: define the hypothesis, build the pipeline, run the model, analyze the result, and decide what to do next. Are comfortable working across research, product, infrastructure, data, evals, and safety boundaries, and can communicate clearly with each group. Like building load‑bearing systems and processes when that is what the team needs, even if the work is not glamorous. Want to train and ship the models that make agents genuinely useful for developers, enterprises, researchers, and everyday users. About the Company the company is an AI research and deployment company dedicated to ensuring that general‑purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. Equal Opportunity Employment We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see the company’s affirmative action and equal employment opportunity policy statement. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. Privacy Notice the company Global Applicant Privacy Policy #J-18808-Ljbffr United States Digital Space LLC
- OpenAI is seeking a Context Researcher on the Agent Post-Training team to scale compute spent on context and to... ...model runs and ship enhancements to frontier products used by developers and... ...You will design experiments, build evals and data pipelines, and contribute to...Training
$380k
Research Scientist - Multimodal Agent, Consumer Devices | OpenAI Careers Research Scientist... .... We work at the frontier of multimodal AI, helping... ...team to work on RLHF and post‑training for personalized, multimodal... ...clean experiments, reliable evals, and decision‑useful...TrainingWork at officeImmediate startRelocation package$180k - $225k
...transformation through frontier AI systems that... ...production AI agents that automate... ...architectures, and research papers emerge every... ...in high-stakes environments.Collaborate with... ...displayed on each job posting reflects the... ...relevant education or training. Scale employees...TrainingFull time$58 - $63 per hour
About the Role The Agents team investigates how to build, align, and scale frontier AI systems that can tackle complex, multi... ...infrastructure—from designing post-training methods for agentic behavior to... ...metrics fall short. As a research intern, you will work on problems...TrainingHourly payInternship$190k - $270k
...TeamThe Databricks AI Research organization is pushing the frontier of next-generation... ...the models and agents that unlock it. Our... ...stack, from model training to advanced multi-agent... ...pillars include post-training enhancements... ...of specialized RL environments. As a member of this...TrainingLocal areaWorldwide- ...building long-horizon multi-agent systems and pushing the boundaries of AI research, this is the kind of role... ...novel problems at the frontier of what's currently possible... ...-horizon reasoning, LLM post‑training for non‑myopic objectives, environment and feedback design....Training
$218.4k - $273k
...accelerating the abundance of frontier data to pave the road... ...the ACE team The Agent Capabilities & Environments (ACE) team, part of Scale’s Research organization, brings... ...on each job posting reflects the minimum... ...relevant education or training. Scale employees in eligible...TrainingFull time$216k - $270k
Scale Labs, Research Scientist - Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral... ...so by elements of their environment; Design and build exploits... ...prototypes. Experience with post-training and RL techniques such as...TrainingFull time- Traverse is a research data lab building reinforcement learning environments for frontier AI labs. As a Research Scientist, you will design and build RL environments that... ...and domain knowledge, turning that into training signals that actually work. We value rigorous...Training
$264.8k - $331k
...About the General Agents TeamThe General... ...intersection of frontier agent development... ...of real customer environments.You will:Design and... ...spaces, balancing research-driven approaches... ...on each job posting reflects the minimum... ...relevant education or training. Scale employees...TrainingFull time$250k - $300k
...models at leading research labs and enterprises... ...solutions for frontier AI development: Enterprise... ...high-quality training data at scale... ...Join Us High-Impact Environment : We operate like... ...evaluate autonomous agent capabilities. Design... ..., and blog posts. What You Bring Ph...TrainingWork at officeFlexible hours2 days per week- Polymath Labs is hiring a Founding Member of Technical Staff - Research to push the frontier of autonomous agents. You will work on long-horizon evaluation, agent post-training, and environment design, wearing multiple hats from building benchmarks to running rigorous...Training
$259.2k - $324k
...accelerating the abundance of frontier data to pave the road... ...the ACE team The Agent Capabilities & Environments (ACE) team, part of Scale’s Research organization, brings... ...on each job posting reflects the minimum... ...relevant education or training. Scale employees in eligible...TrainingFull time$252k - $315k
...systems in production environments.This is a Management role... ...AI models and agents within customer security... ...current on the latest AI/ML research and tools, bringing... ...displayed on each job posting reflects the minimum and... ...relevant education or training. Scale employees in...TrainingFull time- ...to-end software agents. We're the... ...Role MissionPost-training is the critical... ...role blends deep research and hands-on engineering... ...across evals and production... ...systems through post-training,... ...moving research environments where priorities... ...to operate at frontier scale from day...TrainingShift work
- United States Digital Space LLC is looking for a Context Researcher to scale compute on context in the Agent Post-Training team. This role involves collaborating with researchers and engineers on model training and improving product interfaces. Your responsibilities will...Training
- Goaly is seeking an AI Researcher to lead research on agentic AI, focusing on training specialized models and building orchestration stacks. You'll design experiments, publish results, and advance the field. Ideal candidates have a Ph.D. or Master's in relevant fields,...Training
$264.8k - $331k
...from Meta, we are doubling down on building out state of the art post-training algorithms to reach the performance necessary for complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are working...TrainingFull time$25k
...agencies to safeguard the public through information sharing, training, research, and technology. Duties & Responsibilities This is an Open... ...Announcement (OCA) to fill multiple Criminal Investigator (Special Agent) vacancies for the Homeland Security Task Force. Eligible...TrainingWork at officeLocal areaImmediate startRelocation package- ...DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering... ...overview As an AI Researcher, Agent, you will own the full... ...stacks and simulated environments that make large‑scale... ...from data curation through post‑training optimization....TrainingFull timeWork at office
$151.5k - $222.2k
...technologies and capabilities. Frontier AI is a purpose‑built... ...models, multi‑agent systems, and robotics... .... You will design the environments, rewards, and domain‑specific... .... Responsibilities Research & Innovation: Partner... ...scientific signal. Post‑train domain models (SFT, DPO...TrainingFull timeFlexible hours$250k
Research Scientist / Engineer - Multimodal Agent SF Bay Area, CA • Remote, International • London, UK | Research Remote • Hybrid Full-time About Luma AI Luma... ...models, we believe that AI needs to be jointly trained over all signal modalities - text, video, audio, images...TrainingFull timeRemote workWorldwide$216k - $270k
...Software Engineer, Agent Oversight About... ...focused on pushing the frontier of what agentic... ...paired with applied ML research, design, and... ...tradeoffs between offline evals and live customer... ...to own model training, modeling strategy... ...on each job posting reflects the minimum...TrainingFull time$248.4k - $310.5k
...focused on pushing the frontier of what agentic... ...paired with applied ML research, design, and evaluation... ...Engineering Manager for Agent Oversight, you'll lead... ...displayed on each job posting reflects the minimum and... ...relevant education or training. Scale employees in eligible...TrainingFull time$205.6k - $257k
...advancements in AI, including frontier model training, enterprise adoption,... ...vertical within our Agents Data & Reinforcement Learning Environments team. In this role,... ...with a sense for AI research and current agent capabilities... ...displayed on each job posting reflects the minimum...TrainingFull time- ...San Francisco is looking for a Member of Technical Staff - Research to advance autonomous agent capabilities. The ideal candidate will engage in core research areas, including agent evaluation and environment design, while also contributing to building benchmarks and writing...
- ...the Team API Agents builds the shared... ...that turn OpenAI’s frontier models into... ...software engineering, research, finance, healthcare... ..., execution environments, identity and permissions... ...than model training: success comes from... ...believe this job posting is non-compliant,...TrainingFull timeInternship
- Scale is seeking a Machine Learning Research Engineer, Agents - Enterprise GenAI to advance state-of-the-art Agent RL training for enterprise datasets. You will train models, iterate... ...with LLMs in production, experience with post-training methods (RLHF/RLVR, PPO/GRPO), and...Training
- ...About the Team OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Agent Harness Bridge team sits... ...effectively in training environments. This role is ideal... ...you believe this job posting is non-compliant, please...TrainingFull time
$110k - $120k
...built on a diverse culture of trained, talented, and mission-... ...protective services program with a frontier technology client with... ...full-time Executive Protection Agents with high integrity, sound judgment... ...maintain a safe and secure environment. Operate in a low-profile, low...TrainingHourly payDaily paidFull timeTemporary workLocal areaRelocation packageFlexible hoursShift workNight shiftWeekend work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Agent Post-Training, Frontier Evals and Environments Research. Be the first to apply!
- fbi agent San Francisco, CA
- booking agent San Francisco, CA
- tsa agent San Francisco, CA
- state farm agent San Francisco, CA
- import export agent San Francisco, CA
- remote chat agent San Francisco, CA
- agent San Francisco, CA
- executive protection agent San Francisco, CA
- right of way agent San Francisco, CA
- showing agent San Francisco, CA


