Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Practitioner for AI Model Evaluation

$60 - $70 per hour

SaidGig

Role Overview

Help shape the safety, quality, and alignment of frontier AI models by evaluating responses to complex, policy-sensitive, and ambiguous topics. This role contributes structured assessments and feedback that improve model behavior across high-impact domains.

Key Responsibilities

  • Assess AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
  • Review content related to misinformation, political persuasion, self-harm, violence, cybersecurity, biosecurity, and other sensitive subjects.
  • Apply and refine evaluation rubrics used for RLHF, SFT, and AI safety benchmarking.
  • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
  • Provide structured feedback to strengthen model alignment and safety performance.
  • Partner with AI researchers and safety teams on ongoing evaluation initiatives.

Qualifications

  • Bachelor''s degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.
  • At least 5 years of professional experience in AI safety, trust and safety, journalism, public policy, scientific research, security, or a related field.
  • Excellent written English, critical-thinking, and analytical-reasoning skills.
  • Ability to evaluate nuanced, policy-sensitive scenarios consistently.

Preferred Qualifications

  • Experience in AI safety, RLHF, SFT, trust and safety, or AI evaluation.
  • Familiarity with safety policies, content moderation, or developing evaluation rubrics.
  • Experience reviewing complex, high-risk, or ambiguous content.

Work Terms

  • Remote, hourly engagement.

Compensation

  • $60 to $70 per hour.

Opportunity

  • Contribute to frontier AI systems used by millions worldwide.
  • Work on real-world safety evaluations spanning nuanced, high-impact domains.
  • Collaborate with AI researchers, engineers, and safety teams.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Safety Practitioner for AI Model Evaluation in United States vacancy
  • $60 - $70 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Jack Dorsey . Position: AI Safety Practitioner Type: Contract...  ...Role Responsibilities Evaluate AI-generated responses for...  ...structured feedback to improve model alignment and safety performance... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    3 days ago
  • Xcede is looking for a Member of Technical Staff focused on AI Safety to lead red-teaming efforts and ensure the robustness of next-generation AI systems. The selected candidate will design scalable safety frameworks, partner with researchers to define production safety... 
    Suggested

    Xcede

    San Francisco, CA
    3 days ago
  • $218.5k - $288k

     ...Scientist specializing in Small Language Models and AI Training, you will lead research and...  ...language models.Design, implement, and evaluate model training experiments to improve performance...  ...for real-world applications.Ensure AI safety, fairness, and alignment principles are... 
    Suggested
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    2 days ago
  • $60 - $90 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 hour ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines...  ...functional, performance, robustness, and safety metrics, including LLM-judge–based... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    9 hours ago
  • $207k - $285k

    OpenAI is seeking a Technical Program Manager in San Francisco to lead initiatives that ensure the safety and robustness of its AI models. The role involves collaborating with diverse teams to turn risks into actionable plans. Ideal candidates will have experience in technical... 

    OpenAI

    San Francisco, CA
    2 days ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Hourly pay
    Remote work
    Flexible hours

    Prolific

    Jacksonville, FL
    2 days ago
  • $224k - $356.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts...  ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    4 days ago
  • $65 per hour

    Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience... 
    Hourly pay
    Self employment
    Work from home
    Flexible hours

    Prolific

    Las Vegas, NV
    4 days ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Weekday work

    Mercor Inc

    San Francisco, CA
    5 days ago
  • $40 per hour

    A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows full... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Oklahoma City, OK
    4 days ago
  • $150 per hour

     ...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $14 - $42 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Hindi music scene and detailed quality standards. Key Responsibilities Assess AI-generated music and rate it against established quality criteria. Compare AI-generated... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    8 days ago
  •  ...Role Overview Work with a leading AI lab to evaluate outputs from generative music models in German and English. This role focuses on listening, scoring, and annotating AI-generated music and lyrics across genres, using music production and audio engineering vocabulary... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    10 hours per week

    SaidGig

    United States
    a month ago
  • $100 per hour

     ...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas...  ...focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $17 - $54 per hour

     ...Role Overview Evaluate AI-generated music and lyrics across a wide range of genres, applying detailed quality standards in both French and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics for... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    8 days ago
  • $65 - $90 per hour

     ...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and provide... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    19 days ago
  • $35 - $62 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying Korean music expertise and detailed quality standards in both Korean and English. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    8 days ago
  • $85 - $120 per hour

     ...Apply senior-level legal judgment to help a leading AI research team improve how frontier models handle real-world legal work. You will turn specialized...  ...tasks, model instructions, reference solutions, and evaluation benchmarks. Key Responsibilities Review legal knowledge... 
    Hourly pay
    Full time
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    8 days ago
  • $65 - $90 per hour

     ...Apply deep financial judgment to help improve foundational AI models. In this role, you will create and evaluate finance-focused work that strengthens AI systems'' reasoning, analysis, and decision-making capabilities. Role Overview This is a W-2 employment opportunity... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically...  ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    24 days ago
  • $65 - $90 per hour

     ...Overview Provide expert chemical engineering knowledge to evaluate and improve AI systems, ensuring domain accuracy and practical usefulness....  ...standards used in chemical engineering practice, including process safety and operations. Collaborate with a small group of... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    19 days ago
  • $75 per hour

     ...apply geospatial imaging, survey, and photogrammetric expertise to evaluate AI-generated maps and geospatial content, verify spatial accuracy, and provide clear, structured feedback that improves model outputs. No prior AI experience is required. Key Responsibilities... 
    Hourly pay
    Temporary work
    Part time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $75 per hour

     ...Records Managers apply professional records, archives, and library expertise to evaluate AI-generated outputs, create field-accurate prompts, and provide structured feedback that improves model performance on records-management tasks. Candidates can include Archivists,... 
    Hourly pay
    Temporary work
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $50 - $75 per hour

    A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in... 
    Hourly pay
    Contract work

    Mercor

    San Francisco, CA
    4 days ago
  • Apple Inc. is seeking a Research Scientist/Engineer to design evaluation systems for foundation models powering Apple products. You will work hands‑on across evaluation design, experimentation, and cross‑team collaboration to drive model improvement and product quality... 

    Apple Inc.

    Cupertino, CA
    4 days ago
  • $300k - $320k

     ...role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be...  ...of our AI models. Working closely with our Research, Trust & Safety, Frontier Redteaming, and Policy teams, you will drive high-... 
    Work at office
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    Seattle, WA
    2 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    3 days ago
  • Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts... 
    Remote job

    Dorado

    New York, NY
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Practitioner for AI Model Evaluation. Be the first to apply!