AI Safety Practitioner for Model Evaluation
$60 - $70 per hourSaidGig
Help strengthen the safety, quality, and alignment of frontier AI models by evaluating their responses across complex, policy-sensitive, and ambiguous topics. This role focuses on structured assessment and feedback that improves model behavior in high-impact real-world domains. Key Responsibilities
- Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality.
- Review material involving misinformation, political persuasion, self-harm, violence, cyber topics, biosecurity, and other sensitive areas.
- Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking.
- Identify unsafe outputs, hallucinations, reasoning failures, and policy violations.
- Provide structured feedback to improve model alignment and safety performance.
- Collaborate with AI researchers and safety teams on ongoing evaluation initiatives.
- Bachelor''s degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline.
- At least 5 years of professional experience in AI safety, trust and safety, journalism, public policy, scientific research, security, or a related field.
- Excellent written English, critical-thinking, and analytical-reasoning skills.
- Ability to consistently evaluate nuanced, policy-sensitive scenarios.
- Experience with AI safety, RLHF, SFT, trust and safety, or AI evaluation.
- Familiarity with safety policies, content moderation, or evaluation-rubric development.
- Experience reviewing complex, high-risk, or ambiguous content.
- Remote hourly engagement.
- 60 to 70 per hour.
- Shape the safety and behavior of frontier AI models used by millions of people worldwide.
- Work on challenging safety evaluations across nuanced, high-impact domains.
- Collaborate with AI researchers, engineers, and safety teams.
$70 - $110 per hour
...Role Overview Shape how advanced AI systems reason about real clinical work. In this... ...will partner with an AI research team to evaluate medical knowledge tasks, define high-quality... ...develop benchmarks that measure meaningful model improvement. Key Responsibilities Review...SuggestedHourly payFull timeLive inRelocationRelocation package$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to...SuggestedHourly payRemote work$100 per hour
...Role Overview Apply your finance expertise to help improve AI models across complex financial problem-solving areas, including capital... ...prior AI experience is required. Key Responsibilities Evaluate language models in finance domains where performance needs improvement...SuggestedHourly payContract workRemote work10 hours per weekFlexible hours- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...SuggestedFor contractorsRemote work
$70 - $90 per hour
...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels...SuggestedHourly payRemote work$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours$60 - $90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation: $60–...Contract workSummer workRemote work$100 - $150 per hour
...data scientists who will be considered for future projects evaluating how well AI systems perform real-world data science tasks. Members of this... ...criteria, assess AI or human-produced analyses and models, document decisions in writing, and iterate on evaluations with...Hourly payImmediate startRemote work$100 - $150 per hour
...Role Overview Apply your data science expertise to evaluate AI-generated slides, spreadsheets, and documents for real-world quality and usability. You will assess outputs against professional standards and deliver clear feedback that improves their accuracy and presentation...Hourly payWork at officeRemote work- ...Overview Help assess frontier AI systems at the boundary between legitimate radiological safety work and potentially dangerous misuse. You will apply practitioner judgment to determine when a technical... ..., dual-use, or adversarial. Evaluate AI responses against a defined...For contractorsRemote work
$60 - $80 per hour
...expertise to help develop advanced large language models. In this role, you will bring practical brand, growth, and campaign judgment to AI training data, partnering with research... ...reasoning quality. Develop and improve evaluation guidelines and scoring rubrics for...Hourly payWeekday work$100 - $150 per hour
...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work... ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$100 per hour
...Overview Apply your finance expertise to improve AI-driven financial applications by providing rigorous, real-world analysis, evaluation, and feedback. This remote, part-time... ...annotate complex financial data, reports, and model outputs using detailed rubrics and...Hourly payContract workPart timeFor contractorsRemote work$20 per hour
SupportFinity™ in Maine is seeking an Editorial Proofreader to join their team focused on training AI models. The position involves evaluating AI chatbot outputs and improving model quality through expert editing and writing skills. This flexible role allows you to work...Remote jobHourly payFlexible hours- SupportFinity™ is looking for an Editorial Proofreader to join our team to train AI models. In this role, you will measure AI chatbot progress, evaluate logic, and solve problems to enhance model quality. Applicants should have a strong command of English and experience...Remote jobHourly payFull timePart timeFlexible hours
$20 per hour
SupportFinity™ is seeking an Editorial Proofreader to evaluate AI models and improve their quality through expert writing and editing skills. This role can be part‑time or full‑time, allowing for a flexible schedule and project selection. Applicants must be fluent in English...Remote jobHourly payFull timePart timeFlexible hours- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
$20 per hour
SupportFinity™ is looking for an Editorial Proofreader to join our team for AI model training. In this remote role, you'll evaluate AI chatbots and enhance model quality. Candidates should have fluency in English and strong editing skills. This position can be full‑time...Remote jobHourly payFull timeContract workPart time- ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Remote jobHourly payFlexible hours
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours$62.4k - $72.8k
...Why RoboForce RoboForce is an AI robotics company developing Physical AI–powered... ...scalability. We are looking for a Model Evaluation Operator- AI Robotics (Contractor) to help... ...inconsistencies, failure patterns, safety risks, and unexpected robot behaviors....Hourly payContract workFor contractorsMonday to FridayShift workAfternoon shift- ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research... ...role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and...Temporary workFor contractorsRemote work
$50 - $100 per hour
...Apply advanced software engineering and problem-solving expertise to code generation and model evaluation work that helps improve how next-generation AI systems learn, reason, and perform. This remote contract role centers on real-world coding challenges, codebase improvement...Hourly payContract workFor contractorsRemote work- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically... ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation...Contract workFor contractorsFreelanceRemote work
$85 - $105 per hour
...Help improve how advanced AI systems handle real-world contract work by applying hands-on legal experience to contract drafting, review... ...redlining scenarios. This part-time contractor role focuses on evaluating and refining AI performance on technology-focused commercial...Hourly payContract workPart timeFor contractorsRemote work$60 - $80 per hour
...Overview Apply deep insurance expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to...Hourly payWeekday work$60 - $80 per hour
...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting... ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities...Hourly payContract workRemote work$60 - $90 per hour
...hands-on mechanical engineering judgment to improve how advanced AI models reason through real-world engineering work. You will partner... ..., define standards for correct solutions, and create rigorous evaluations grounded in industry practice. Key Responsibilities...Hourly payFull timeRemote work$110 per hour
...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application... ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models...Hourly payContract workRemote work$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific...Hourly payFull timeFreelanceLive inRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Safety Practitioner for Model Evaluation. Be the first to apply!
- reiki practitioner United States
- assistant practitioner United States
- stretch practitioner United States
- holistic practitioner United States
- practitioner United States
- infection control practitioner United States
- warehouse safety United States
- safety training United States
- safety scientist United States
- patient safety monitor United States

