Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Researcher, Agent Safety, Training and Evaluations

$380k

OpenAI

About the TeamThe Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously.Our work spans three areas:Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks.Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work.Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review).About the RoleWe’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to.This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.In this role, you will:Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity.Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals.Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems.You might thrive in this role if you:Have demonstrated strength in research engineering, ML engineering, quantitative research, or applied model research, with the ability to own ambiguous projects end to end.Bring excellent technical execution across experimentation, data, evaluation, and/or infrastructure, plus strong intuition for modern frontier-model research.Are motivated by agent safety and eager to work on urgent, practical problems even if your prior work was outside safety or alignment.About OpenAIOpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Researcher, Agent Safety, Training and Evaluations in San Francisco, CA vacancy
  • $380k

    About the TeamThe Agent Safety team works to ensure that increasingly...  ....Our work spans three areas:Training: Create training methods, environments...  ...risks.Measurements: Build evaluations and production metrics that...  ...a safety&security minded researcher or engineer who can reason... 
    Training
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    3 days ago
  • $295k

    About the TeamThe Safety Systems team is responsible for various...  ...transparency.The Model Safety Research team aims to fundamentally advance...  ...such as RLHF, adversarial training, robustness, and more....  ...highest safety standards.Actively evaluate and understand the safety of... 
    Training
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    8 days ago
  •  ...Primarily On-site We are seeking an Agent Evaluation Infrastructure Engineer to build the...  ...and supporting infrastructure used to train and assess long-horizon enterprise AI agents...  ...You will work on the engineering and research problems behind realistic agent... 
    Training

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    5 days ago
  • Thinking Machines in San Francisco seeks a researcher to advance internal evaluations and signals for post-training models. You will collaborate with researchers and engineers across the research organization, shaping evaluation creation, usability, and auditing. Your... 
    Training

    Mosaic.tech

    San Francisco, CA
    2 days ago
  • # Researcher, Frontier Cybersecurity RisksOn-siteSan FranciscoAll...  ...is a critical Safety Research team at...  ...that assist humans to agents that can plan, execute...  ...model capabilities.- Evaluate technical trade-offs within...  ...familiar with methods for training and fine-tuning large... 
    Training

    Triwill Group

    San Francisco, CA
    5 days ago
  •  ...that assist humans to agents that can plan, execute...  ...measures, RSI-relevant training interventions, and...  ...Design experiments and evaluations to understand the extent...  ...misaligned, or their safety-relevant capabilities...  ...rigorous hypothesis-driven research and turning our... 
    Training

    OpenAI

    San Francisco, CA
    2 days ago
  •  ..., and respond appropriately. The Chat and Multimodal Safety team is responsible for ensuring that OpenAI’s increasingly...  ...safely across these experiences. We develop the research, training methods, and evaluations needed to make these experiences safe. Our work sits at... 
    Training
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  • Scale Labs is seeking a Research Scientist focused on Agent Robustness to advance safe and aligned AI agents. You will contribute to evaluating risks, building testing harnesses, and prototyping...  ...emphasizes collaboration, post-training techniques, and publishing results... 
    Training

    United States Digital Space LLC

    San Francisco, CA
    2 days ago
  •  ...Group in San Francisco, CA is seeking an Agent Evaluation Infrastructure Engineer to build...  ...the supporting infrastructure used to train and assess long-horizon enterprise AI agents...  .... You will work on the engineering and research problems behind realistic agent environments... 
    Training

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    3 days ago
  • $151.5k - $222.2k

     ...with the highest possible safety margins. We are...  ...foundation models, multi-agent systems, and robotics to...  ...them. Responsibilities: Research & Innovation Partner with...  ...signal. Post-train domain models (SFT, DPO...  ...scout emerging trends. Evaluate external vendors, open-... 
    Training
    Full time
    Flexible hours

    Initial Therapeutics, Inc.

    San Francisco, CA
    1 day ago
  • $23 - $23.5 per hour

     ...Warehouse AgentDo you enjoy working in a fast-paced, safety-obsessed aviation environment? As a Warehouse Agent, you will be essential to increase operational...  ...operating in a safe and responsible manner.Complete all training when required by company, airport governing... 
    Training
    Full time
    Part time
    Work experience placement
    Work at office
    Immediate start
    Flexible hours
    Night shift

    AGI

    San Francisco, CA
    2 days ago
  • $150k - $250k

     ...global social organizations.We research and deploy technologies that...  ...own construction processes, evaluate them, and evolve. They draw from...  ...(e.g. graph-based planners, agent orchestration frameworks,...  ...systems using models rather than training or fine-tuning them. Ideal... 
    Training
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    4 days ago
  • $52 per hour

     ...professionals passionate about security and safety. Join our innovative team committed to...  ...key roles, such as executive protection agents, intelligence analysts, armed security operatives...  ...Successfully completing required annual training and maintaining operational readiness... 
    Training
    Hourly pay
    Work at office
    Local area
    Immediate start
    Flexible hours
    Shift work
    Night shift
    Weekend work

    Enhanced Protection Services

    San Francisco, CA
    2 days ago
  • $296.3k - $374.8k

     ..., and observability portfolio. Our research team is working on a fundamental problem...  ...Build simulation environments for training and evaluating models and agents before they operate on production...  ...to evaluate the reliability and safety of learned policies before deployment... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    2 days ago
  • $216.3k - $280.8k

     ...we are leading frontier AI research across Cisco. Our mission is...  ...systems, scalable training algorithms, evaluation science, inference optimization...  ...agentic AI workflows, multi-agent frameworks, and autonomous...  ...security-focused use cases, safety, and robust machine learning... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    2 days ago
  • $116.2k - $146.7k

    Associate Professional Researcher - Andino LabThe Andino Laboratory at UCSF is seeking an Associate...  ..., pathogenesis, and the development and evaluation of next-generation live-attenuated...  ...Associate Professional Researcher will train and mentor new members of the laboratory... 
    Training

    University of California, San Francisco

    San Francisco, CA
    4 days ago
  • About the Company We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze...  ...agents plan, decompose problems, choose what to try next, evaluate their own outputs, and recover from mistakes. This is a... 
    Shift work

    MakerMaker.AI

    San Francisco, CA
    5 days ago
  • $293k - $405k

     ...This position requires a deep technical understanding of security measures and active engagement with stakeholders to design and evaluate defense systems. Ideal candidates are proficient in modern software engineering, capable of creating prototypes, and have experience... 

    OpenAI

    San Francisco, CA
    5 days ago
  • $96.7k - $126.4k

    Assistant Professional Researcher - Okada labThe Okada lab in the Department...  ...guidance, technical training, and oversight of laboratory...  ...laboratory members in laboratory safety, cell culture techniques,...  ...needed. Provide ongoing feedback, evaluate trainee progress, and... 
    Training
    Traineeship
    Flexible hours
    Afternoon shift

    University of California, San Francisco

    San Francisco, CA
    1 day ago
  • $172.5k - $260.1k

     ...#1 AI CRM, where humans with agents drive customer success together...  ...are the future of Salesforce.Research & Insights (R&I) is a core...  ...Conduct strategic, generative, and evaluative research using the...  ...class enablement and on-demand training with Trailhead.comExposure to... 
    Training
    Full time

    Salesforce

    San Francisco, CA
    6 days ago
  • $118k - $309.8k

    Educational Research Scientist Department of Surgery and the Center for Faculty Educators...  ...to advance education scholarship and evaluate the effectiveness of educational interventions...  ...with international surgical skills training programs. This is a 1.0 FTE position. -... 
    Training
    Full time
    Work at office

    University of California, San Francisco

    San Francisco, CA
    4 days ago
  •  ...Tilde Research is a moonshot AI lab advancing mechanistic interpretability, new architectures, and pretraining science....  ...use those insights to make them better. You'll work on training, analyzing, and evaluating cutting-edge models, collaborating closely with a team of... 
    Training

    Tilde Research

    San Francisco, CA
    3 days ago
  • $150k - $300k

     ...Join to apply for the Applied Researcher role at Variant Overview Variant is hiring an applied researcher to join our team in...  ...with creativity and taste. This role involves designing, training, and evaluating deep neural networks and large language models (LLMs); prototyping... 
    Training
    Full time

    Variant

    San Francisco, CA
    5 days ago
  • $250k - $350k

     ...boundaries of what's possible in LLM post-training. If you love training models, exploring...  ..., running experiments, and turning research insights into products that ship, we'd love...  ...everything end-to-end: distillation, training, evaluation, and planet-scale hosting. We are a... 
    Training
    Work at office

    Inference

    San Francisco, CA
    3 days ago
  • $218.7k - $249.6k

     ...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with Academia to building...  ...models through all phases of development, from design through training, evaluation, validation, and implementation. Engage in high impact applied... 
    Training
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    5 days ago
  •  ...include molecular dynamics, post-training biomolecular models, and probe-based model evaluation. Close the loop between in...  ...or are eager to develop, strong research judgment in biomolecular modeling...  ...history building production-scale agent platforms. Pay & Benefits... 
    Training
    H1b
    Visa sponsorship

    CAPABLE.org

    San Francisco, CA
    3 days ago
  • $262.5k - $299.6k

    Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating trustworthy and reliable AI systems, changing banking...  ...through all phases of development, from design through training, evaluation, validation, and implementation. Engage in high impact... 
    Training
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    5 days ago
  • $262.5k - $299.6k

    Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One, we are creating trustworthy and reliable...  ...through all phases of development, from design through training, evaluation, validation, and implementation. Engage in high impact applied... 
    Training
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    5 days ago
  • $147k - $210k

     ...projects by defining key research questions.Design, implement, and evaluate experiments to provide...  ...reinforcement learning, multimodal agents, computational social...  ...of experience with training, evaluating, and...  ...scientific discovery, ensuring safety and ethics are always... 
    Training

    Google

    San Francisco, CA
    4 days ago
  • $200k - $280k

     ...algorithms, architectures, engines) and post-training / RL systems. We build and operate the...  ...and/or training stack. Have a solid research foundation in your area(s) of depth: Track...  ...make large-scale rollout collection and evaluation cheaper. Use these pipelines to train,... 
    Training
    Full time

    Together

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Researcher, Agent Safety, Training and Evaluations. Be the first to apply!