Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Practitioner for Model Evaluation

$60 - $70 per hour

SaidGig

Help strengthen the safety, quality, and alignment of frontier AI models by evaluating their responses across complex, policy-sensitive, and ambiguous topics. This role focuses on structured assessment and feedback that improves model behavior in high-impact real-world domains. Key Responsibilities

  • Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
  • Review material involving misinformation, political persuasion, self-harm, violence, cyber topics, biosecurity, and other sensitive areas.
  • Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.
  • Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
  • Provide structured feedback to improve model alignment and safety performance.
  • Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.
Qualifications
  • Bachelor''s degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.
  • At least 5 years of professional experience in AI safety, trust and safety, journalism, public policy, scientific research, security, or a related field.
  • Excellent written English, critical-thinking, and analytical-reasoning skills.
  • Ability to consistently evaluate nuanced, policy-sensitive scenarios.
Preferred Qualifications
  • Experience with AI safety, RLHF, SFT, trust and safety, or AI evaluation.
  • Familiarity with safety policies, content moderation, or evaluation-rubric development.
  • Experience reviewing complex, high-risk, or ambiguous content.
Work Terms
  • Remote hourly engagement.
Compensation
  • 60 to 70 per hour.
What You Will Contribute
  • Shape the safety and behavior of frontier AI models used by millions of people worldwide.
  • Work on challenging safety evaluations across nuanced, high-impact domains.
  • Collaborate with AI researchers, engineers, and safety teams.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Safety Practitioner for Model Evaluation in United States vacancy
  • $70 - $110 per hour

     ...Role Overview Shape how advanced AI systems reason about real clinical work. In this...  ...will partner with an AI research team to evaluate medical knowledge tasks, define high-quality...  ...develop benchmarks that measure meaningful model improvement. Key Responsibilities Review... 
    Suggested
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    Remote
    a month ago
  • $100 per hour

     ...Role Overview Apply your finance expertise to help improve AI models across complex financial problem-solving areas, including capital...  ...prior AI experience is required. Key Responsibilities Evaluate language models in finance domains where performance needs improvement... 
    Suggested
    Hourly pay
    Contract work
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets... 
    Suggested
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  • $70 - $90 per hour

     ...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    Remote
    a month ago
  • $65 per hour

    Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience... 
    Hourly pay
    Self employment
    Work from home
    Flexible hours

    Prolific

    Las Vegas, NV
    3 days ago
  • $60 - $90 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation: $60–... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    4 days ago
  • $100 - $150 per hour

     ...data scientists who will be considered for future projects evaluating how well AI systems perform real-world data science tasks. Members of this...  ...criteria, assess AI or human-produced analyses and models, document decisions in writing, and iterate on evaluations with... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $100 - $150 per hour

     ...Role Overview Apply your data science expertise to evaluate AI-generated slides, spreadsheets, and documents for real-world quality and usability. You will assess outputs against professional standards and deliver clear feedback that improves their accuracy and presentation... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    11 days ago
  •  ...Overview Help assess frontier AI systems at the boundary between legitimate radiological safety work and potentially dangerous misuse. You will apply practitioner judgment to determine when a technical...  ..., dual-use, or adversarial. Evaluate AI responses against a defined... 
    For contractors
    Remote work

    SaidGig

    United States
    19 days ago
  • $60 - $80 per hour

     ...expertise to help develop advanced large language models. In this role, you will bring practical brand, growth, and campaign judgment to AI training data, partnering with research...  ...reasoning quality. Develop and improve evaluation guidelines and scoring rubrics for... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    more than 2 months ago
  • $100 - $150 per hour

     ...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work...  ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $100 per hour

     ...Overview Apply your finance expertise to improve AI-driven financial applications by providing rigorous, real-world analysis, evaluation, and feedback. This remote, part-time...  ...annotate complex financial data, reports, and model outputs using detailed rubrics and... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $20 per hour

    SupportFinity™ in Maine is seeking an Editorial Proofreader to join their team focused on training AI models. The position involves evaluating AI chatbot outputs and improving model quality through expert editing and writing skills. This flexible role allows you to work... 
    Remote job
    Hourly pay
    Flexible hours

    SupportFinity™

    Montgomery, AL
    2 days ago
  • SupportFinity™ is looking for an Editorial Proofreader to join our team to train AI models. In this role, you will measure AI chatbot progress, evaluate logic, and solve problems to enhance model quality. Applicants should have a strong command of English and experience... 
    Remote job
    Hourly pay
    Full time
    Part time
    Flexible hours

    SupportFinity™

    Columbia, SC
    2 days ago
  • $20 per hour

    SupportFinity™ is seeking an Editorial Proofreader to evaluate AI models and improve their quality through expert writing and editing skills. This role can be part‑time or full‑time, allowing for a flexible schedule and project selection. Applicants must be fluent in English... 
    Remote job
    Hourly pay
    Full time
    Part time
    Flexible hours

    SupportFinity™

    Sioux Falls, SD
    2 days ago
  • Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote job
    Flexible hours

    Prolific

    Charlotte, NC
    1 day ago
  • $20 per hour

    SupportFinity™ is looking for an Editorial Proofreader to join our team for AI model training. In this remote role, you'll evaluate AI chatbots and enhance model quality. Candidates should have fluency in English and strong editing skills. This position can be full‑time... 
    Remote job
    Hourly pay
    Full time
    Contract work
    Part time

    SupportFinity™

    New York, NY
    1 day ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    1 day ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago
  • $62.4k - $72.8k

     ...Why RoboForce RoboForce is an AI robotics company developing Physical AI–powered...  ...scalability. We are looking for a Model Evaluation Operator- AI Robotics (Contractor) to help...  ...inconsistencies, failure patterns, safety risks, and unexpected robot behaviors.... 
    Hourly pay
    Contract work
    For contractors
    Monday to Friday
    Shift work
    Afternoon shift

    RoboForce

    Milpitas, CA
    3 days ago
  •  ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research...  ...role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and... 
    Temporary work
    For contractors
    Remote work

    micro1

    Remote
    15 days ago
  • $50 - $100 per hour

     ...Apply advanced software engineering and problem-solving expertise to code generation and model evaluation work that helps improve how next-generation AI systems learn, reason, and perform. This remote contract role centers on real-world coding challenges, codebase improvement... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    19 days ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically...  ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $85 - $105 per hour

     ...Help improve how advanced AI systems handle real-world contract work by applying hands-on legal experience to contract drafting, review...  ...redlining scenarios. This part-time contractor role focuses on evaluating and refining AI performance on technology-focused commercial... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    2 days ago
  • $60 - $80 per hour

     ...Overview Apply deep insurance expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    more than 2 months ago
  • $60 - $80 per hour

     ...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting...  ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $60 - $90 per hour

     ...hands-on mechanical engineering judgment to improve how advanced AI models reason through real-world engineering work. You will partner...  ..., define standards for correct solutions, and create rigorous evaluations grounded in industry practice. Key Responsibilities... 
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    16 days ago
  • $110 per hour

     ...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application...  ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Practitioner for Model Evaluation. Be the first to apply!