AI Safety Practitioner for Model Evaluation
$60 - $70 per hourSaidGig
Help strengthen the safety, quality, and alignment of frontier AI models by evaluating their responses across complex, policy-sensitive, and ambiguous topics. This role focuses on structured assessment and feedback that improves model behavior in high-impact real-world domains. Key Responsibilities
- Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
- Review material involving misinformation, political persuasion, self-harm, violence, cyber topics, biosecurity, and other sensitive areas.
- Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.
- Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
- Provide structured feedback to improve model alignment and safety performance.
- Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.
- Bachelor''s degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.
- At least 5 years of professional experience in AI safety, trust and safety, journalism, public policy, scientific research, security, or a related field.
- Excellent written English, critical-thinking, and analytical-reasoning skills.
- Ability to consistently evaluate nuanced, policy-sensitive scenarios.
- Experience with AI safety, RLHF, SFT, trust and safety, or AI evaluation.
- Familiarity with safety policies, content moderation, or evaluation-rubric development.
- Experience reviewing complex, high-risk, or ambiguous content.
- Remote hourly engagement.
- 60 to 70 per hour.
- Shape the safety and behavior of frontier AI models used by millions of people worldwide.
- Work on challenging safety evaluations across nuanced, high-impact domains.
- Collaborate with AI researchers, engineers, and safety teams.
$70 - $110 per hour
...Role Overview Shape how advanced AI systems reason about real clinical work. In this... ...will partner with an AI research team to evaluate medical knowledge tasks, define high-quality... ...develop benchmarks that measure meaningful model improvement. Key Responsibilities Review...SuggestedHourly payFull timeLive inRelocationRelocation package$60 - $90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation...SuggestedFull timeContract workSummer workRemote work$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...SuggestedHourly payRemote workWork from homeFlexible hours$60 per hour
...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing...SuggestedHourly payRemote workWork from homeFlexible hours$70 - $80 per hour
...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports...SuggestedHourly payContract workRemote work$60 - $80 per hour
...expertise to help develop advanced large language models. In this role, you will bring practical brand, growth, and campaign judgment to AI training data, partnering with research... ...reasoning quality. Develop and improve evaluation guidelines and scoring rubrics for...Hourly payWeekday work- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$100 - $150 per hour
...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work... ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$70 - $90 per hour
...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels...Hourly payRemote work$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to...Hourly payRemote work$100 per hour
...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas... ...focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours- ...underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for... ...Box. Deepgram’s voice-native foundation models are accessed through cloud APIs or as... ...looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible...Full time
- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically... ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation...Contract workFor contractorsFreelanceRemote work
$100 - $150 per hour
...data scientists who will be considered for future projects evaluating how well AI systems perform real-world data science tasks. Members of this... ...criteria, assess AI or human-produced analyses and models, document decisions in writing, and iterate on evaluations with...Hourly payImmediate startRemote work$60 - $80 per hour
...Overview Apply deep insurance expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to...Hourly payWeekday work$36 - $72 per hour
...educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role... ...that probe where a model's quality or safety behavior breaks down Identify...Hourly payFull timeMonday to FridayFlexible hours$60 - $80 per hour
...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting... ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities...Hourly payContract workRemote work- ...’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and... ...About the role We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do....Full time
$100 per hour
...Apply consulting expertise to evaluate and improve AI-generated business content for a customer-facing project. Your judgment will help AI systems... ...executive summaries. Develop and refine large language model prompts using structured problem-solving and analytical rigor...Hourly payPart timeFor contractorsRemote work$110 per hour
...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application... ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models...Hourly payContract workRemote work$65 - $105 per hour
...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world software development and engineered... ...program management team to define high-quality engineering work, evaluate model performance, and turn expert practice into clear...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific...Hourly payFull timeFreelanceLive inRelocationRelocation package$60 - $90 per hour
...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong...Hourly payWork at officeRemote work$90 - $175 per hour
...Role Overview Apply software quality assurance expertise to evaluate technical AI outputs and help improve how next-generation AI systems... ...through human data annotation, labeling, RLHF, AI response or model evaluation, or rubric-based grading. Software QA experience...Hourly payContract workRemote work- ...Apply your dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical, remote opportunity with no direct patient care, focused on ensuring AI-generated medical outputs reflect real...Hourly payRemote work
$100 - $150 per hour
...Apply senior finance expertise to improve how frontier AI systems reason through real-world financial work. You will partner closely... ...with an AI research team to define high-quality finance tasks, evaluate model performance, and translate professional judgment into rigorous...Hourly payFull timeLive inRelocationRelocation package$60 - $90 per hour
...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,... ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities...Hourly payPart time$20 - $36 per hour
...Role Overview Evaluate AI-generated music in Slovak and English by listening across genres and providing detailed ratings against quality... ...assessments and labels will help improve generative music models. Key Responsibilities Compare pairs of AI-generated songs...Hourly payImmediate startRemote workFlexible hours$60 - $80 per hour
...Role Overview Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative solutions grounded in real-world practice, and judge model outputs to improve model correctness...Hourly payWeekday work$100 - $150 per hour
...matter expertise to a GenAI research team, creating authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role centers on hands-on legal judgment, translating professional...Hourly payFull timeFreelanceInternshipLive inLocal areaRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Safety Practitioner for Model Evaluation. Be the first to apply!
- holistic practitioner United States
- infection control practitioner United States
- stretch practitioner United States
- reiki practitioner United States
- practitioner United States
- assistant practitioner United States
- oilfield safety United States
- safety internship United States
- safety valve technician United States
- safety associate United States




