Researcher, Agent Safety, Training and Evaluations
$380kOpenAI
About the TeamThe Agent Safety team works to ensure that increasingly capable AI agents act safely, exercise sound judgment, and remain aligned with user intent. Our mission is to reduce the probability of severe unintended outcomes from increasingly capable AI agents while preserving their ability to act effectively and autonomously.Our work spans three areas:Training: Create training methods, environments and data that teach agents to make better decisions in consequential situations. We turn real-world failures into training signals that prevent similar incidents, and identify precursor behaviors and mitigations to address emerging risks.Measurements: Build evaluations and production metrics that identify emerging risks and measure whether our interventions work.Oversight: Develop oversight and system mitigation mechanisms that reduce harmful actions while preserving useful autonomy (for example future versions of auto-review).About the RoleWe’re looking for strong executors with excellent judgment, comfort with ambiguity, and an understanding of frontier model research. You don’t need prior safety or alignment experience, we also welcome people that recently realized that alignment and safety is a critical area to contribute to.This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.In this role, you will:Train and evaluate frontier models to reduce harmful or misaligned agent actions, forming clear hypotheses and executing independently through ambiguity.Mine incidents and build scalable measurement, data-processing, and evaluation systems that turn real failures into repeatable safety signals.Collaborate closely with post-training, capabilities, oversight, and pre-training partners to ship research-backed mitigations into large-scale training and agent systems.You might thrive in this role if you:Have demonstrated strength in research engineering, ML engineering, quantitative research, or applied model research, with the ability to own ambiguous projects end to end.Bring excellent technical execution across experimentation, data, evaluation, and/or infrastructure, plus strong intuition for modern frontier-model research.Are motivated by agent safety and eager to work on urgent, practical problems even if your prior work was outside safety or alignment.About OpenAIOpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an
$380k
About the TeamThe Agent Safety team works to ensure that increasingly... ....Our work spans three areas:Training: Create training methods, environments... ...risks.Measurements: Build evaluations and production metrics that... ...a safety&security minded researcher or engineer who can reason...TrainingWork at officeRelocation package$295k
About the TeamThe Safety Systems team is responsible for various... ...transparency.The Model Safety Research team aims to fundamentally advance... ...such as RLHF, adversarial training, robustness, and more.... ...highest safety standards.Actively evaluate and understand the safety of...TrainingWork at officeLocal areaFlexible hours- ...Primarily On-site We are seeking an Agent Evaluation Infrastructure Engineer to build the... ...and supporting infrastructure used to train and assess long-horizon enterprise AI agents... ...You will work on the engineering and research problems behind realistic agent...Training
- Thinking Machines in San Francisco seeks a researcher to advance internal evaluations and signals for post-training models. You will collaborate with researchers and engineers across the research organization, shaping evaluation creation, usability, and auditing. Your...Training
- # Researcher, Frontier Cybersecurity RisksOn-siteSan FranciscoAll... ...is a critical Safety Research team at... ...that assist humans to agents that can plan, execute... ...model capabilities.- Evaluate technical trade-offs within... ...familiar with methods for training and fine-tuning large...Training
- ...that assist humans to agents that can plan, execute... ...measures, RSI-relevant training interventions, and... ...Design experiments and evaluations to understand the extent... ...misaligned, or their safety-relevant capabilities... ...rigorous hypothesis-driven research and turning our...Training
- ..., and respond appropriately. The Chat and Multimodal Safety team is responsible for ensuring that OpenAI’s increasingly... ...safely across these experiences. We develop the research, training methods, and evaluations needed to make these experiences safe. Our work sits at...TrainingWork at officeRelocation package
- Scale Labs is seeking a Research Scientist focused on Agent Robustness to advance safe and aligned AI agents. You will contribute to evaluating risks, building testing harnesses, and prototyping... ...emphasizes collaboration, post-training techniques, and publishing results...Training
- ...Group in San Francisco, CA is seeking an Agent Evaluation Infrastructure Engineer to build... ...the supporting infrastructure used to train and assess long-horizon enterprise AI agents... .... You will work on the engineering and research problems behind realistic agent environments...Training
$151.5k - $222.2k
...with the highest possible safety margins. We are... ...foundation models, multi-agent systems, and robotics to... ...them. Responsibilities: Research & Innovation Partner with... ...signal. Post-train domain models (SFT, DPO... ...scout emerging trends. Evaluate external vendors, open-...TrainingFull timeFlexible hours$23 - $23.5 per hour
...Warehouse AgentDo you enjoy working in a fast-paced, safety-obsessed aviation environment? As a Warehouse Agent, you will be essential to increase operational... ...operating in a safe and responsible manner.Complete all training when required by company, airport governing...TrainingFull timePart timeWork experience placementWork at officeImmediate startFlexible hoursNight shift$150k - $250k
...global social organizations.We research and deploy technologies that... ...own construction processes, evaluate them, and evolve. They draw from... ...(e.g. graph-based planners, agent orchestration frameworks,... ...systems using models rather than training or fine-tuning them. Ideal...TrainingWork at office3 days per week$52 per hour
...professionals passionate about security and safety. Join our innovative team committed to... ...key roles, such as executive protection agents, intelligence analysts, armed security operatives... ...Successfully completing required annual training and maintaining operational readiness...TrainingHourly payWork at officeLocal areaImmediate startFlexible hoursShift workNight shiftWeekend work$296.3k - $374.8k
..., and observability portfolio. Our research team is working on a fundamental problem... ...Build simulation environments for training and evaluating models and agents before they operate on production... ...to evaluate the reliability and safety of learned policies before deployment...TrainingFull timeTemporary workLocal areaFlexible hours$216.3k - $280.8k
...we are leading frontier AI research across Cisco. Our mission is... ...systems, scalable training algorithms, evaluation science, inference optimization... ...agentic AI workflows, multi-agent frameworks, and autonomous... ...security-focused use cases, safety, and robust machine learning...TrainingFull timeTemporary workLocal areaFlexible hours$116.2k - $146.7k
Associate Professional Researcher - Andino LabThe Andino Laboratory at UCSF is seeking an Associate... ..., pathogenesis, and the development and evaluation of next-generation live-attenuated... ...Associate Professional Researcher will train and mentor new members of the laboratory...Training- About the Company We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze... ...agents plan, decompose problems, choose what to try next, evaluate their own outputs, and recover from mistakes. This is a...Shift work
$293k - $405k
...This position requires a deep technical understanding of security measures and active engagement with stakeholders to design and evaluate defense systems. Ideal candidates are proficient in modern software engineering, capable of creating prototypes, and have experience...$96.7k - $126.4k
Assistant Professional Researcher - Okada labThe Okada lab in the Department... ...guidance, technical training, and oversight of laboratory... ...laboratory members in laboratory safety, cell culture techniques,... ...needed. Provide ongoing feedback, evaluate trainee progress, and...TrainingTraineeshipFlexible hoursAfternoon shift$172.5k - $260.1k
...#1 AI CRM, where humans with agents drive customer success together... ...are the future of Salesforce.Research & Insights (R&I) is a core... ...Conduct strategic, generative, and evaluative research using the... ...class enablement and on-demand training with Trailhead.comExposure to...TrainingFull time$118k - $309.8k
Educational Research Scientist Department of Surgery and the Center for Faculty Educators... ...to advance education scholarship and evaluate the effectiveness of educational interventions... ...with international surgical skills training programs. This is a 1.0 FTE position. -...TrainingFull timeWork at office- ...Tilde Research is a moonshot AI lab advancing mechanistic interpretability, new architectures, and pretraining science.... ...use those insights to make them better. You'll work on training, analyzing, and evaluating cutting-edge models, collaborating closely with a team of...Training
$150k - $300k
...Join to apply for the Applied Researcher role at Variant Overview Variant is hiring an applied researcher to join our team in... ...with creativity and taste. This role involves designing, training, and evaluating deep neural networks and large language models (LLMs); prototyping...TrainingFull time$250k - $350k
...boundaries of what's possible in LLM post-training. If you love training models, exploring... ..., running experiments, and turning research insights into products that ship, we'd love... ...everything end-to-end: distillation, training, evaluation, and planet-scale hosting. We are a...TrainingWork at office$218.7k - $249.6k
...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with Academia to building... ...models through all phases of development, from design through training, evaluation, validation, and implementation. Engage in high impact applied...TrainingFull timePart timeLocal areaFlexible hours- ...include molecular dynamics, post-training biomolecular models, and probe-based model evaluation. Close the loop between in... ...or are eager to develop, strong research judgment in biomolecular modeling... ...history building production-scale agent platforms. Pay & Benefits...TrainingH1bVisa sponsorship
$262.5k - $299.6k
Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating trustworthy and reliable AI systems, changing banking... ...through all phases of development, from design through training, evaluation, validation, and implementation. Engage in high impact...TrainingFull timePart timeLocal areaFlexible hours$262.5k - $299.6k
Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One, we are creating trustworthy and reliable... ...through all phases of development, from design through training, evaluation, validation, and implementation. Engage in high impact applied...TrainingFull timePart timeLocal areaFlexible hours$147k - $210k
...projects by defining key research questions.Design, implement, and evaluate experiments to provide... ...reinforcement learning, multimodal agents, computational social... ...of experience with training, evaluating, and... ...scientific discovery, ensuring safety and ethics are always...Training$200k - $280k
...algorithms, architectures, engines) and post-training / RL systems. We build and operate the... ...and/or training stack. Have a solid research foundation in your area(s) of depth: Track... ...make large-scale rollout collection and evaluation cheaper. Use these pipelines to train,...TrainingFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Researcher, Agent Safety, Training and Evaluations. Be the first to apply!
- field researcher San Francisco, CA
- product researcher San Francisco, CA
- security researcher San Francisco, CA
- lead researcher San Francisco, CA
- data collection researcher San Francisco, CA
- machine learning researcher San Francisco, CA
- court researcher San Francisco, CA
- researcher San Francisco, CA
- postdoctoral researcher cosmetic science San Francisco, CA
- senior researcher San Francisco, CA


