AI Safety Practitioner - Expert Evaluator
$60 - $70 per hourMercor
Job Description
Job Description
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: AI Safety Practitioner
Type: Contract
Compensation: $60–$70/hour
Location: Remote
Role Responsibilities
- Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
- Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains.
- Apply and refine evaluation rubrics for RLHF , SFT , and AI safety benchmarking .
- Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
- Provide structured feedback to improve model alignment and safety performance.
- Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.
Qualifications
Must-Have
- Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.
- 5+ years of professional experience in AI Safety , Trust & Safety, journalism, public policy, scientific research, security, or a related field.
- Excellent written English , critical thinking, and analytical reasoning skills.
- Ability to consistently evaluate nuanced and policy-sensitive scenarios.
Preferred
- Experience with AI Safety , RLHF , SFT , Trust & Safety, or AI evaluation.
- Familiarity with safety policies, content moderation, or evaluation rubric development.
- Experience reviewing complex, high-risk, or ambiguous content.
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to: View email address on ziprecruiter.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
- Obsidian is seeking expert Evaluators in FP&A / corporate finance to assess AI-generated work products for accuracy and quality. This role entails deep expertise to grade outputs and provide structured feedback. Candidates should have at least 5 years of relevant experience...SuggestedRemote jobHourly payWork at office
- Obsidian is seeking expert Evaluators in Biology/environmental science to review and assess AI-generated work products for accuracy and quality. In this remote, hourly role, you will leverage your expertise to provide feedback on documents and presentations, ensuring they...SuggestedRemote jobHourly pay
- ...Welo Data is seeking Data Labeling Associates in California to evaluate AI outputs and ensure cultural context and safety in Arabic datasets. This role requires professional-level proficiency in Portuguese (Brazil), a bachelor's degree, and at least 2 years of experience...Suggested
$90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...genuinely correct replications from those that merely look correct. Evaluate responsive behavior and semantic quality, ensuring proper use of...SuggestedContract workSummer workLocal areaRemote work- Synthires is offering a part-time role for PhD-level Chemistry experts to contribute to AI safety and evaluation projects. The work involves applying scientific expertise to understand and improve how AI systems handle specialized chemistry topics, with training provided...SuggestedRemote jobPart time
- Obsidian seeks experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models across grey-area topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and...
- ...Francisco is hiring Data Labeling Associates for Project Perseus. This role focuses on evaluating Arabic AI systems, requiring professional proficiency in Portuguese and experience in AI safety. Responsibilities include assessing AI outputs, identifying bias, and...Full time
$70 - $84 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ..., and Jack Dorsey . Position: AI Safety Red Teamer Type: Contract Compensation... ...hallucinations, and policy failures. Evaluate model robustness across misinformation, cyber...Contract workSummer workRemote work- We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area"...
- ...Mercor, we believe the safest AI is the one that’s already been... ...for this project - human data experts who probe AI models with adversarial... ...customer AI systems Evaluation coverage expands: more scenarios... ...production Mercor customers trust the safety of their AI because you’ve...Remote work
$29 - $45 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Portuguese (global) Type: Contract Compensation:...Contract workSummer workRemote work- YO AI Labs is seeking a Pharmacovigilance Expert to contribute drug safety expertise to a healthcare AI project. You will review pharmacovigilance documentation, safety... ...The role emphasizes data quality, analytical evaluation of DSURs/PSURs, and compliance with ICH E2F,...Remote jobFor contractorsFlexible hours
- ...Data is seeking Data Labeling Associates in San Francisco to evaluate AI systems focused on Arabic language nuances. The role includes... ...like gourmet dining and comprehensive medical coverage while contributing to innovative AI safety solutions. #J-18808-Ljbffr Welo Data
$50 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Jack Dorsey . Position: Spanish (Spain) Audio Generalist Evaluator Expert Type: Contract Compensation: $50/hour Location:...Contract workSummer workRemote work$15 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Jack Dorsey . Position: Music & Lyrics Expert - Malayalam Type: Contract Compensation... ...+ hours/week Role Responsibilities Evaluate AI-generated music across various genres...Contract workSummer workImmediate startRemote workFlexible hours$80 - $120 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Position: Pricing / ROI / revenue economics Evaluator Type: Contract Compensation: $... ...performance . Collaborate with subject matter experts to ensure consistency and domain...Contract workSummer workWork at officeRemote work- About the roleWe're building a high-quality evaluation dataset for CNC manufacturing and are looking for experienced CNC machinists to help author and validate grading rubrics for CNC machining work. You'll bring real production-floor judgment to determine whether a machining...
$120 - $175 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Jack Dorsey . Position: CNC Machining Expert Type: Contract Compensation: $1... ...~ Strong production judgment to assess safety and feasibility on real machines. ~ Written...Contract workSummer workRemote work$100 - $150 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Dorsey . Position: Sales Engineering Expert Type: Contract Compensation: $... ...objections, supporting proofs of concept, and evaluating product fit and technical risk. ~ Hands...Contract workSummer workRemote work- 1. Role Overview Mercor is partnering with a leading AI research organization to engage experienced sales engineers for a project focused on evaluating how well AI systems perform real-world technical sales work. Rather than producing deliverables yourself, you will define...
- Obsidian is looking for expert Evaluators in Finance operations/audit support to review AI-generated work products for accuracy and quality. This remote hourly position requires a minimum of 5 years in finance and fluency in English. Your role will involve evaluating outputs...Remote jobHourly payWork at office
- Mercor is seeking experienced musicians to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated lyrics across a wide range of genres and rate them against detailed quality standards, working in Malayalam and English. Required...
$80 - $120 per hour
Mercor is seeking a User/Customer Research and Feedback Synthesis Evaluator to evaluate AI-generated artifacts using specific quality rubrics. The role demands deep subject-matter expertise in user research and involves providing feedback to improve AI performance. The...Remote jobContract workWork at officeFlexible hours- Obsidian is seeking a Spanish Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project. You will handle transcription, annotation, and evaluation tasks to help train and benchmark advanced language models. The ideal candidate should have...Part time10 hours per week
- Obsidian is hiring expert Evaluators in Investment analysis / valuation / credit to review AI-generated work products for accuracy and quality. This remote, hourly position requires deep subject-matter expertise and professional fluency in English to provide structured...Hourly payWork at officeRemote work
$80 - $120 per hour
Mercor is looking for a Biology / environmental science Evaluator to assess AI-generated artifacts based on quality rubrics. This position requires evaluating documents for errors and collaborating with AI teams to improve model performance. The ideal candidate should...Remote jobHourly payContract workWork at office- Obsidian is hiring expert Evaluators in Real estate, hospitality, and events to review AI-generated work for accuracy, rigor, and domain quality. This remote position requires deep expertise and involves grading outputs like documents and presentations. Applicants must...Remote jobWork at office
$80 - $120 per hour
Mercor is seeking a Media / journalism / communications Evaluator to evaluate AI-generated artifacts and provide structured feedback. Ideal candidates will have over 5 years of experience and must be fluent in English. This remote role requires proficiency in Microsoft...Remote jobHourly payWork at office$50 - $75 per hour
A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in...Hourly payContract work$93 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...identification, categorization, and resolution of claim denials. Evaluate AI-generated appeal letters, denial root cause analyses, and...Contract workSummer workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Safety Practitioner - Expert Evaluator. Be the first to apply!
- practitioner San Francisco, CA
- holistic practitioner San Francisco, CA
- infection control practitioner San Francisco, CA
- technology expert San Francisco, CA
- subject matter expert San Francisco, CA
- fulfillment expert San Francisco, CA
- guest service support expert San Francisco, CA
- work from home web search evaluator San Francisco, CA
- social media evaluator San Francisco, CA
- evaluator San Francisco, CA



