Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Practitioner - Expert Evaluator

Mercor

We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. Responsibilities Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality. Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains. Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking. Identify unsafe outputs, hallucinations, reasoning failures, and policy violations. Provide structured feedback to improve model alignment and safety performance. Collaborate with AI researchers and safety teams on ongoing evaluation initiatives. Required Qualifications Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline. 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field. Excellent written English, critical thinking, and analytical reasoning skills. Ability to consistently evaluate nuanced and policy-sensitive scenarios. Preferred Qualifications Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation. Familiarity with safety policies, content moderation, or evaluation rubric development. Experience reviewing complex, high-risk, or ambiguous content. Why Join? Shape the safety and behaviour of frontier AI models used by millions worldwide. Work on challenging, real-world safety evaluations across nuanced and high-impact domains. Collaborate with leading AI researchers, engineers, and safety teams. #J-18808-Ljbffr Mercor

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Safety Practitioner - Expert Evaluator in San Francisco, CA vacancy
  • $60 - $70 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ..., and Jack Dorsey . Position: AI Safety Practitioner Type: Contract Compensation:...  ...Remote Role Responsibilities Evaluate AI-generated responses for safety, factual... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    5 days ago
  • Obsidian is seeking expert Evaluators in FP&A / corporate finance to assess AI-generated work products for accuracy and quality. This role entails deep expertise to grade outputs and provide structured feedback. Candidates should have at least 5 years of relevant experience... 
    Suggested
    Remote job
    Hourly pay
    Work at office

    Obsidian

    San Francisco, CA
    4 days ago
  • Obsidian is seeking expert Evaluators in Biology/environmental science to review and assess AI-generated work products for accuracy and quality. In this remote, hourly role, you will leverage your expertise to provide feedback on documents and presentations, ensuring they... 
    Suggested
    Remote job
    Hourly pay

    Obsidian

    San Francisco, CA
    5 days ago
  • $90 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...genuinely correct replications from those that merely look correct. Evaluate responsive behavior and semantic quality, ensuring proper use of... 
    Suggested
    Contract work
    Summer work
    Local area
    Remote work

    Mercor

    San Francisco, CA
    8 days ago
  • $80 - $120 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Jack Dorsey . Position: BI dashboards / performance reporting Evaluator Type: Contract Compensation: $80–$120/hour... 
    Suggested
    Contract work
    Summer work
    Work at office
    Remote work

    Mercor

    San Francisco, CA
    14 days ago
  • Welo Data is seeking Data Labeling Associates in California to evaluate AI outputs and ensure cultural context and safety in Arabic datasets. This role requires professional-level proficiency in Portuguese (Brazil), a bachelor's degree, and at least 2 years of experience... 

    Welo Data

    San Francisco, CA
    1 day ago
  • Synthires is offering a part-time role for PhD-level Chemistry experts to contribute to AI safety and evaluation projects. The work involves applying scientific expertise to understand and improve how AI systems handle specialized chemistry topics, with training provided... 
    Remote job
    Part time

    Synthires

    San Francisco, CA
    3 days ago
  • Obsidian seeks experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models across grey-area topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and... 

    Obsidian

    San Francisco, CA
    2 days ago
  •  ...Francisco is hiring Data Labeling Associates for Project Perseus. This role focuses on evaluating Arabic AI systems, requiring professional proficiency in Portuguese and experience in AI safety. Responsibilities include assessing AI outputs, identifying bias, and... 
    Full time

    Welo Data

    San Francisco, CA
    5 days ago
  • We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area"... 

    Obsidian

    San Francisco, CA
    1 day ago
  •  ...Mercor, we believe the safest AI is the one that’s already been...  ...for this project - human data experts who probe AI models with adversarial...  ...customer AI systems Evaluation coverage expands: more scenarios...  ...production Mercor customers trust the safety of their AI because you’ve... 
    Remote work

    Obsidian

    San Francisco, CA
    4 days ago
  • $70 - $84 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Summers , and Jack Dorsey . Position AI Safety Red Teamer Type Contract Compensation $70...  ...behaviors, hallucinations, and policy failures. Evaluate model robustness across misinformation,... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    2 days ago
  • $29 - $45 per hour

    Mercor connects elite creative and technical talent with leading AI research labs. The AI Safety Experts — English & Portuguese (global) contract role is remote, offering $29-$45/hour. You will red team conversational AI models, generate high-quality human data, and document... 
    Remote job
    Contract work

    Mercor

    San Francisco, CA
    4 days ago
  • $29 - $45 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Portuguese (global) Type: Contract Compensation:... 
    Contract work
    Summer work
    Remote work

    Mercor Inc

    San Francisco, CA
    4 days ago
  • YO AI Labs is seeking a Pharmacovigilance Expert to contribute drug safety expertise to a healthcare AI project. You will review pharmacovigilance documentation, safety...  ...The role emphasizes data quality, analytical evaluation of DSURs/PSURs, and compliance with ICH E2F,... 
    Remote job
    For contractors
    Flexible hours

    YO AI Labs

    San Francisco, CA
    4 days ago
  • $20 - $22 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Assamese Type: Contract Compensation: $2... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    7 days ago
  • $48 - $62 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Dutch Type: Contract Compensation: $48–$... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    9 days ago
  •  ...Data is seeking Data Labeling Associates in San Francisco to evaluate AI systems focused on Arabic language nuances. The role includes...  ...like gourmet dining and comprehensive medical coverage while contributing to innovative AI safety solutions. #J-18808-Ljbffr Welo Data

    Welo Data

    San Francisco, CA
    3 days ago
  • Mercor is seeking experienced AI Safety Practitioners to assess the safety, quality, and alignment of frontier AI models across complex, policy-sensitive topics. You will evaluate AI-generated responses, apply safety policies, and provide structured feedback to improve... 

    Mercor

    San Francisco, CA
    3 days ago
  •  ...intraoperative needs under physician orders. You will document perioperative care, coordinate with surgeons and anesthesiologists, and perform certain procedures outside the OR according to KP policy, maintaining patient safety and compliance. #J-18808-Ljbffr Kaiser Permanente

    Kaiser Permanente

    San Francisco, CA
    2 days ago
  • $120 - $175 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Larry Summers , and Jack Dorsey . Position: CNC Machining Expert Type: Contract Compensation: $120–$175/hour Location... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    11 days ago
  • Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Bring a background in songwriting... 

    Obsidian

    San Francisco, CA
    1 day ago
  • About the roleWe're building a high-quality evaluation dataset for CNC manufacturing and are looking for experienced CNC machinists to help author and validate grading rubrics for CNC machining work. You'll bring real production-floor judgment to determine whether a machining... 

    Obsidian

    San Francisco, CA
    3 days ago
  • $50 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Jack Dorsey . Position: Spanish (Spain) Audio Generalist Evaluator Expert Type: Contract Compensation: $50/hour Location:... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    23 days ago
  • $100 - $150 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Jack Dorsey . Position: B2B Sales Expert Type: Contract Compensation: $...  ...asynchronously to meet deadlines while improving evaluation processes. Qualifications Must-Have... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    5 days ago
  • $80 - $120 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Position: Pricing / ROI / revenue economics Evaluator Type: Contract Compensation: $...  ...performance . Collaborate with subject matter experts to ensure consistency and domain... 
    Contract work
    Summer work
    Work at office
    Remote work

    Mercor

    San Francisco, CA
    15 days ago
  • $15 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Jack Dorsey . Position: Music & Lyrics Expert - Malayalam Type: Contract Compensation...  ...+ hours/week Role Responsibilities Evaluate AI-generated music across various genres... 
    Contract work
    Summer work
    Immediate start
    Remote work
    Flexible hours

    Mercor

    San Francisco, CA
    11 days ago
  • 1. Role Overview Mercor is partnering with a leading AI research organization to engage experienced sales engineers for a project focused on evaluating how well AI systems perform real-world technical sales work. Rather than producing deliverables yourself, you will define... 

    Mercor

    San Francisco, CA
    2 days ago
  • $80 - $120 per hour

    Mercor is looking for a Biology / environmental science Evaluator to assess AI-generated artifacts based on quality rubrics. This position requires evaluating documents for errors and collaborating with AI teams to improve model performance. The ideal candidate should... 
    Remote job
    Hourly pay
    Contract work
    Work at office

    Mercor

    San Francisco, CA
    4 days ago
  • Obsidian is looking for expert Evaluators in Finance operations/audit support to review AI-generated work products for accuracy and quality. This remote hourly position requires a minimum of 5 years in finance and fluency in English. Your role will involve evaluating outputs... 
    Remote job
    Hourly pay
    Work at office

    Obsidian

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Practitioner - Expert Evaluator. Be the first to apply!