Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Evaluator: Frontiers & Policy Alignment

Mercor

Mercor is seeking experienced AI Safety Practitioners to evaluate frontier AI models for safety, alignment, and policy compliance across nuanced topics. You will assess AI-generated responses for quality, identify unsafe outputs, and apply robust evaluation rubrics for RLHF and SFT. The role involves detailed feedback and collaboration with researchers and safety engineers to push forward safe AI behavior. This is a hands-on, impactful position in a dynamic environment, designed for serious #J-18808-Ljbffr Mercor

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Safety Evaluator: Frontiers & Policy Alignment in San Francisco, CA vacancy
  • Obsidian seeks experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models across grey-area topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and... 
    Policy

    Obsidian

    San Francisco, CA
    5 days ago
  • $60 - $70 per hour

     ...technical talent with leading AI research labs....  ...Dorsey . Position: AI Safety Practitioner Type:...  ...Role Responsibilities Evaluate AI-generated responses...  ..., factual accuracy, policy compliance, and overall...  ...feedback to improve model alignment and safety performance.... 
    Policy
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    11 days ago
  •  ...is seeking a Product Marketing Manager, Research & Safety to shape how the world understands frontier AI and our approach to responsible development. You will...  ...enterprise audiences. You will work with Research, Safety, Policy, Communications, and Product teams to lead go-to-... 
    Policy

    OpenAI

    San Francisco, CA
    5 days ago
  • TikTok is seeking a Technical AI Policy Researcher within the Trust & Safety division to advance responsible frontier AI work across multiple teams in San Francisco. You will...  ..., and contribute to policy design, evaluation workflows, and governance activities. The... 
    Policy

    TikTok

    San Francisco, CA
    2 days ago
  •  ...Welo Data is seeking Data Labeling Associates in California to evaluate AI outputs and ensure cultural context and safety in Arabic datasets. This role requires professional-level proficiency in Portuguese (Brazil), a bachelor's degree, and at least 2 years of experience... 
    Suggested

    Welo Data

    San Francisco, CA
    3 days ago
  •  ...Preparedness is a critical Safety Research team at...  ...on mitigating AI threats to global...  ...capabilities of frontier AI systems. Mitigation...  ...safeguards, alignment tools, and security...  ...model capabilities. Evaluate technical trade-...  ...fine‑tuning, and policy optimization. Excel... 
    Policy

    United States Digital Space LLC

    San Francisco, CA
    4 days ago
  •  ...Preparedness is a critical Safety Research team at...  ...on mitigating AI threats to global...  ...capabilities of frontier AI systems. Mitigation...  ...safeguards, alignment tools, and security...  ...preparedness capability evaluations—designing new...  ...Employment Opportunity Policy Statement.... 
    Policy
    Permanent employment
    Temporary work

    United States Digital Space LLC

    San Francisco, CA
    4 days ago
  • $275k - $375k

     ...deployment of new products as we advance frontier, safe AI technology. You will work closely with...  .... About Anthropic Anthropic is an AI safety and research company that’s working to...  ...team has experience across ML, physics, policy, business and product. Responsibilities... 
    Policy
    Work at office
    Home office
    Visa sponsorship
    Relocation package
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...Francisco is hiring Data Labeling Associates for Project Perseus. This role focuses on evaluating Arabic AI systems, requiring professional proficiency in Portuguese and experience in AI safety. Responsibilities include assessing AI outputs, identifying bias, and... 
    Full time

    Welo Data

    San Francisco, CA
    3 days ago
  • Federation of American Scientists invites interns for the Policy Entrepreneurship Fellowship focused on frontier AI safety. The position runs Sept 30, 2026 - Feb 28, 2027 with a hybrid arrangement in Washington, DC. Fellows work part-time as FAS affiliates, developing... 
    Policy
    Part time

    EBRC

    Emeryville, CA
    3 days ago
  •  ...focused on Agent Robustness to advance safe and aligned AI agents. You will contribute to evaluating risks, building testing harnesses, and...  ...techniques, and publishing results to shape policy and industry practices in frontier AI. #J-18808-Ljbffr United States Digital... 
    Policy

    United States Digital Space LLC

    San Francisco, CA
    5 days ago
  • $108k - $208.8k

    Responsibilities The Trust & Safety (T&S) Responsible AI Policy team's mission is to ensure the development...  ...responsible development and deployment of our frontier AI models across multiple businesses...  ..., and drive end-to-end policy to evaluate workflows for your domain areas.... 
    Policy
    Temporary work
    Local area

    TikTok

    San Francisco, CA
    2 days ago
  •  ...are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through...  ...model weaknesses, and evaluate AI behavior across complex...  ..., hallucinations, and policy failures. Evaluate...  ...researchers to improve model alignment, robustness, and... 
    Policy

    Obsidian

    San Francisco, CA
    4 days ago
  • $150k

    Join Amazon's Frontier AI & Robotics team as a Member of Technical Staff, this Technical Program...  ...and partner organizations, ensuring alignment on commitments, timelines, and...  ...federal, state, and local laws and Company policies. Criminal history may have a direct, adverse... 
    Policy
    Local area
    Day shift

    Amazon

    San Francisco, CA
    2 days ago
  • $164.7k - $339.08k

     ...Possible.At Pinterest, AI isn't just a...  ...Manager for the GenAI Safety team within Trust...  ...with engineering, policy, data science, and...  ...'ll work at the frontier of responsible AI...  ...teaming exercises and evaluation frameworks before...  ...teams to define, align, and ship AI... 
    Policy
    Work at office
    Local area
    Remote work
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $179.4k - $224.25k

    About ScaleAt Scale AI, our mission is to accelerate...  ...the abundance of frontier data to pave the road...  ...upon our prior model evaluation work with enterprise...  ..., model evaluation, safety, and alignment. The data we are...  ...USDPLEASE NOTE: Our policy requires a 90-day waiting... 
    Policy
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  •  ...chain‑of‑thought of frontier reasoning models...  ...methods, and improving alignment. We were the...  ...practical additional safety mechanism, and...  ...training, alignment evaluations, monitoring, and...  ...increasingly capable AI systems more...  ...employment opportunity policy statement.... 
    Policy
    Work at office
    Relocation package
    3 days per week

    United States Digital Space LLC

    San Francisco, CA
    2 days ago
  • $230k - $325k

    About the TeamOpenAI’s Safety teams work to ensure our products are...  ..., trusted, and resilient as frontier AI systems scale globally. We...  ...intersection of product, safety, policy, and research.About the...  ...on your background and team alignment, you may work on areas such as... 
    Policy
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • $205.6k - $257k

     ...been the leading AI data foundry, helping...  ...in AI, including frontier model training,...  ...environments, and evaluations frontier labs use...  ...generation and regression safety, secure code...  ...Security, Legal, and Policy to make these...  ...ability to drive alignment across cross-functional... 
    Policy
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $365k

     ...interpretable, and steerable AI systems. We want AI to...  ..., engineers, policy experts, and business...  ...and post-training to alignment, interpretability, and safety, each operating at the frontier of AI development. As...  ...exposure to training, evaluation, or large-scale distributed... 
    Policy
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    4 days ago
  • We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve... 
    Policy
    Worldwide

    Mercor

    San Francisco, CA
    3 days ago
  •  ...important part of the Safety Systems org at OpenAI,...  ...Preparedness Framework. Frontier AI models have the...  ...capability assessment, evaluations, and internal red teaming...  ...actionable product and policy changes.This position...  ...interest in AI safety, alignment, and catastrophic risk... 
    Policy
    Shift work

    OpenAI

    San Francisco, CA
    2 days ago
  •  ...Data is seeking Data Labeling Associates in San Francisco to evaluate AI systems focused on Arabic language nuances. The role includes...  ...like gourmet dining and comprehensive medical coverage while contributing to innovative AI safety solutions. #J-18808-Ljbffr Welo Data

    Welo Data

    San Francisco, CA
    1 day ago
  • $150k

    Join our Frontier AI & Robotics team to lead the hardware integration...  ...design, ensuring quality and alignment with engineering priorities....  ...inventory, vendor coordination, and safety/regulatory compliance. -...  ..., and local laws and Company policies. Criminal history may have a... 
    Policy
    Local area

    Amazon

    San Francisco, CA
    2 days ago
  • $293k - $405k

     ...Preparedness is a critical Safety Research team at...  ...on mitigating AI threats to global...  ...capabilities of frontier AI systems. Mitigation...  ...safeguards, alignment tools, and security...  ...defenses with AI agent evaluations and penetration...  ...Opportunity Policy Statement. Background... 
    Policy

    Neura Market

    San Francisco, CA
    4 days ago
  • $70 - $84 per hour

     ...talent with leading AI research labs. Headquartered...  .... Position: AI Safety Red Teamer Type:...  ...to stress-test frontier AI models . Identify...  ...hallucinations, and policy failures. Evaluate model robustness across...  ...to improve model alignment, robustness, and safety... 
    Policy
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    10 days ago
  • $216.2k - $270.25k

    About ScaleAt Scale AI, our mission is to accelerate...  ...the abundance of frontier data to pave the road...  ...upon our prior model evaluation work with enterprise...  ..., model evaluation, safety, and alignment. The data we are...  ...USDPLEASE NOTE: Our policy requires a 90-day waiting... 
    Policy
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $246.8k - $333.9k

     ...you will lead the AWS account team for a frontier AI model lab and one of Amazon's largest...  ...closely with our partner organization to align on GTM goal execution- Build and maintain...  ...federal, state, and local laws and Company policies. Criminal history may have a direct,... 
    Policy
    Contract work
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    San Francisco, CA
    4 days ago
  • $350k

     ...interpretable, and steerable AI systems. We want...  ..., engineers, policy experts, and...  ...Research Engineer on Alignment Science, you'll...  ...experimental research on AI safety, with a focus on...  ...-Tuning, and the Frontier Red Team. Our...  ...with third-party evaluators.Safeguards Research... 
    Policy
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    2 days ago
  • $290k - $365k

     ...interpretable, and steerable AI systems. We want AI to...  ..., engineers, policy experts, and business leaders...  ...breakthrough in AI safety research and every interaction...  ...clusters for training frontier models, production...  ...product development Drive alignment on priorities and... 
    Policy
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Evaluator: Frontiers & Policy Alignment. Be the first to apply!