Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Practitioner for Model Evaluation

$60 - $70 per hour

SaidGig

Help strengthen the safety, quality, and alignment of frontier AI models by evaluating their responses across complex, policy-sensitive, and ambiguous topics. This role focuses on structured assessment and feedback that improves model behavior in high-impact real-world domains. Key Responsibilities

  • Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
  • Review material involving misinformation, political persuasion, self-harm, violence, cyber topics, biosecurity, and other sensitive areas.
  • Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.
  • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
  • Provide structured feedback to improve model alignment and safety performance.
  • Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.
Qualifications
  • Bachelor''s degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.
  • At least 5 years of professional experience in AI safety, trust and safety, journalism, public policy, scientific research, security, or a related field.
  • Excellent written English, critical-thinking, and analytical-reasoning skills.
  • Ability to consistently evaluate nuanced, policy-sensitive scenarios.
Preferred Qualifications
  • Experience with AI safety, RLHF, SFT, trust and safety, or AI evaluation.
  • Familiarity with safety policies, content moderation, or evaluation-rubric development.
  • Experience reviewing complex, high-risk, or ambiguous content.
Work Terms
  • Remote hourly engagement.
Compensation
  • 60 to 70 per hour.
What You Will Contribute
  • Shape the safety and behavior of frontier AI models used by millions of people worldwide.
  • Work on challenging safety evaluations across nuanced, high-impact domains.
  • Collaborate with AI researchers, engineers, and safety teams.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Safety Practitioner for Model Evaluation in United States vacancy
  • $70 - $110 per hour

     ...Role Overview Shape how advanced AI systems reason about real clinical work. In this...  ...will partner with an AI research team to evaluate medical knowledge tasks, define high-quality...  ...develop benchmarks that measure meaningful model improvement. Key Responsibilities Review... 
    Suggested
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    22 days ago
  • $60 - $90 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    18 hours ago
  • $60 per hour

     ...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Suggested
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago
  • $60 per hour

     ...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing... 
    Suggested
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    2 days ago
  • $70 - $80 per hour

     ...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports... 
    Suggested
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $60 - $80 per hour

     ...expertise to help develop advanced large language models. In this role, you will bring practical brand, growth, and campaign judgment to AI training data, partnering with research...  ...reasoning quality. Develop and improve evaluation guidelines and scoring rubrics for... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  •  ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets... 
    For contractors
    Remote work

    SaidGig

    United States
    16 days ago
  • $100 - $150 per hour

     ...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work...  ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago
  • $70 - $90 per hour

     ...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    18 days ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    18 days ago
  • $100 per hour

     ...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas...  ...focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for...  ...Box. Deepgram’s voice-native foundation models are accessed through cloud APIs or as...  ...looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible... 
    Full time

    Deepgram

    United States
    18 hours ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically...  ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  • $100 - $150 per hour

     ...data scientists who will be considered for future projects evaluating how well AI systems perform real-world data science tasks. Members of this...  ...criteria, assess AI or human-produced analyses and models, document decisions in writing, and iterate on evaluations with... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  • $60 - $80 per hour

     ...Overview Apply deep insurance expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    more than 2 months ago
  • $36 - $72 per hour

     ...educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role...  ...that probe where a model's quality or safety behavior breaks down Identify... 
    Hourly pay
    Full time
    Monday to Friday
    Flexible hours

    Handshake

    Seattle, WA
    6 days ago
  • $60 - $80 per hour

     ...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting...  ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    more than 2 months ago
  •  ...’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and...  ...About the role We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do.... 
    Full time

    Anthropic

    New York, NY
    3 days ago
  • $100 per hour

     ...Apply consulting expertise to evaluate and improve AI-generated business content for a customer-facing project. Your judgment will help AI systems...  ...executive summaries. Develop and refine large language model prompts using structured problem-solving and analytical rigor... 
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    22 days ago
  • $110 per hour

     ...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application...  ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $65 - $105 per hour

     ...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world software development and engineered...  ...program management team to define high-quality engineering work, evaluate model performance, and turn expert practice into clear... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    22 days ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    22 days ago
  • $60 - $90 per hour

     ...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    12 days ago
  • $90 - $175 per hour

     ...Role Overview Apply software quality assurance expertise to evaluate technical AI outputs and help improve how next-generation AI systems...  ...through human data annotation, labeling, RLHF, AI response or model evaluation, or rubric-based grading. Software QA experience... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    4 days ago
  •  ...Apply your dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical, remote opportunity with no direct patient care, focused on ensuring AI-generated medical outputs reflect real... 
    Hourly pay
    Remote work

    SaidGig

    United States
    3 days ago
  • $100 - $150 per hour

     ...Apply senior finance expertise to improve how frontier AI systems reason through real-world financial work. You will partner closely...  ...with an AI research team to define high-quality finance tasks, evaluate model performance, and translate professional judgment into rigorous... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago
  • $60 - $90 per hour

     ...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,...  ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities... 
    Hourly pay
    Part time

    SaidGig

    Oregon State
    a month ago
  • $20 - $36 per hour

     ...Role Overview Evaluate AI-generated music in Slovak and English by listening across genres and providing detailed ratings against quality...  ...assessments and labels will help improve generative music models. Key Responsibilities Compare pairs of AI-generated songs... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $60 - $80 per hour

     ...Role Overview Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative solutions grounded in real-world practice, and judge model outputs to improve model correctness... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    1 day ago
  • $100 - $150 per hour

     ...matter expertise to a GenAI research team, creating authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role centers on hands-on legal judgment, translating professional... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Local area
    Relocation
    Relocation package

    SaidGig

    California
    7 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Practitioner for Model Evaluation. Be the first to apply!