Behavioral Health Professional for AI Model Evaluation
$45 - $70 per hourSaidGig
Help shape how AI models respond to sensitive everyday conversations involving relationships, family dynamics, emotional wellbeing, personal beliefs and difficult life decisions. You will evaluate model interactions and define standards for responses that are balanced, respectful, neutral and safely bounded, particularly in non-crisis situations where sound judgment matters.
Key Responsibilities- Review user and AI conversations covering relationship and family advice, emotional wellbeing, spiritual or metaphysical questions, and unconventional or unfounded beliefs.
- Assess whether model responses are neutral, appropriate and safe.
- Identify and document concerns such as excessive agreement, taking sides, reinforcing distorted or unfounded beliefs, moralizing, or providing overly clinical or directive advice.
- Create rubrics, guidelines and reference responses for supportive, balanced and appropriately limited replies grounded in counseling and behavioral-health practice.
- Develop test scenarios that assess model behavior in sensitive, non-crisis conversations.
- Work with researchers and fellow experts to maintain consistent, calibrated and well-documented evaluation standards.
- Degree in psychology, counseling, social work, behavioral health, behavioral science, human services or a closely related field, or equivalent mental-health professional experience.
- At least 3 years of professional experience supporting people in a mental health, counseling or social-services setting, such as therapist, counselor, clinical social worker, psychologist, psychiatric nurse, case manager, crisis counselor, peer support specialist or mental health advocate.
- Ability to remain neutral and nonjudgmental across diverse perspectives, relationships, belief systems and worldviews, and clearly explain professional judgments.
- Working knowledge of sycophancy, cognitive distortions, healthy boundaries and client-centered approaches such as motivational interviewing.
- Strong written communication skills and the ability to provide precise, well-structured feedback.
Clinical licensure, including LMFT, LCSW, LPC, LMHC, PsyD or PhD, is valued but not required. Helpful backgrounds include AI safety, applied ethics, trust and safety, content policy, couples or family counseling, spiritual care, religious or alternative belief communities, misinformation or conspiracy-belief psychology, and AI evaluation, annotation or red teaming. Candidates do not need every preferred qualification to apply.
Work Terms- United States-based hourly W-2 employment.
- Part-time commitment of at least 20 hours per week, with the option to work up to 40 hours per week.
- Reliable weekday availability is required.
- Employees may be placed with a leading AI lab as part of its extended workforce.
$45 to $70 per hour.
Equal OpportunityEmployment decisions are made without discrimination based on race, religion, color, national origin, sex, pregnancy or related conditions, sexual orientation, gender identity or expression, age, protected veteran status, disability, genetic information, political views or activity, or any other legally protected characteristic.
$45 - $70 per hour
Gridnaut Recruiting is hiring a remote Behavioral Health Expert, AI Safety and Model Evaluation contractor (pay $45-$70/hr). Contribute to frontier AI research... ...work. Ideal candidates: Behavioral health professionals (therapists, counselors, social workers, psychologists...SuggestedTemporary workPart timeFor contractorsRemote work$70 - $110 per hour
...Overview Shape how advanced AI systems reason about... ...AI research team to evaluate medical knowledge... ...that measure meaningful model improvement. Key Responsibilities... ...outputs for missing behaviors, weak reasoning,... ...Medical Officer. Professional, hands-on experience...SuggestedHourly payFull timeLive inRelocationRelocation package$100 - $150 per hour
...expertise to improve how frontier AI systems reason through real-world... ...high-quality finance tasks, evaluate model performance, and translate professional judgment into rigorous standards.... ...model outputs, identifying missing behaviors, weak reasoning, flawed assumptions...SuggestedHourly payFull timeLive inRelocationRelocation package- Cincinnatus LLC is seeking behavioral health experts to help evaluate how AI models respond in everyday conversations about relationships, wellbeing and beliefs. This part-time role focuses on neutrality, safety and balanced guidance, with 20-40 hours per week and a W-...SuggestedPart time
$110 per hour
...physician talent network supporting AI labs and companies with medical expertise... ...real-world clinical expertise to model development and evaluation. Key Responsibilities Train... ...frontier AI research. Qualifications Professional experience in clinical diagnosis and...SuggestedHourly payContract workRemote work$100 per hour
...finance expertise to improve AI-driven financial applications... ...rigorous, real-world analysis, evaluation, and feedback. This remote,... ...produce accurate, clear, and professionally relevant financial content. Prior... ...financial data, reports, and model outputs using detailed...Hourly payContract workPart timeFor contractorsRemote work$50 per hour
...performance of large language models on real-world finance tasks by evaluating model outputs, designing... ...working directly with AI researchers to shape... ...Responsibilities Evaluate LLM behavior and performance in... ...Minimum 2 years of professional experience in one or more...Hourly payContract workRemote work10 hours per weekFlexible hours$60 - $80 per hour
...help develop advanced large language models by bringing real-world underwriting,... ...claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities... ...Qualifications At least 8 years of dedicated professional insurance experience in underwriting,...Hourly payWeekday work$100 - $150 per hour
...how next-generation AI models perform real financial... ...identifying missing behaviors, weak reasoning, flawed... ...would not survive professional scrutiny. Instruction... ...finance tasks and evaluation sets, and collaborate... ...childbirth, reproductive health decisions, or related...Hourly payFull timeLive inRelocationRelocation package$70 - $100 per hour
...improve how next-generation AI models complete spreadsheet-based work... ...for versatile Excel professionals with experience across multiple... ...human data annotation, or model evaluation is a strong plus.... ...pregnancy, childbirth, reproductive health decisions, related medical conditions...Hourly payWeekday work$60 - $80 per hour
...Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative... ...Qualifications Eight or more years of focused professional experience in retail, such as merchandising, category...Hourly payWeekday work$110 - $150 per hour
...Role Overview Help advance frontier AI models by bringing professional finance judgment to the evaluation, design, and improvement of financial knowledge-work tasks... ...tasks and model outputs, identifying missing behaviors, weak reasoning, flawed assumptions, and responses...Hourly payFull timeLive inRelocationRelocation package$100 per hour
...domain expertise to improve AI-driven financial applications... ...reviewing, annotating, and refining model outputs and prompts. You will... .... Develop, refine, and evaluate prompts related to financial... ...robustness, and adherence to professional financial standards....Hourly payPart timeFor contractorsRemote work- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets,... .... Provide domain feedback that directly shapes model behavior. Required Skills Financial modeling. Valuation....For contractorsRemote work
$60 - $100 per hour
...advance frontier AI systems by bringing... ...judgment into the evaluation, training, and improvement of models used for real insurance... ...missing behaviors, weak reasoning, incorrect... ...would not meet professional standards. Write... ...and casualty, health, reinsurance, pricing...Hourly payFull timeLive inRelocationRelocation package$40 - $90 per hour
...expertise to help train next-generation AI systems through rigorous, real-world medical evaluation work. You will create and assess... ..., resident, physician, or professional with equivalent biomedical... ...or machine-learning systems and model evaluation is helpful but not required...Hourly payContract workRemote work$60 - $80 per hour
...underwriting and claims judgment, and evaluate large language model outputs against structured rubrics to... ...and claims practice. Evaluate AI model outputs against structured rubrics... ...Qualifications Minimum 8 years of dedicated professional experience in insurance, such as...Hourly payWeekday work$110 per hour
...Provide clinical expertise to help train, evaluate, and shape medical AI systems by joining a Physician... ...Responsibilities Train and evaluate AI models in medical and clinical contexts.... ...project needs. Qualifications Professional experience in clinical diagnosis and...Hourly payContract workRemote work$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous life sciences... ...judgment into task design, evaluation, and model improvement. You... ...and model outputs for missing behaviors, weak reasoning, unsupported... ...of large language models in professional work and sound judgment in distinguishing...Hourly payFull timeFreelanceLive inRelocationRelocation package- ...We are seeking an experienced Engineer, AI – AI Evaluation & Model Risk Lead to lead how AI models are evaluated, cleared, monitored, and... ...technical guidance. Required Qualifications ~2+ years of professional experience in relevant AI, machine learning, software...Full timeWork experience placement
$136.44k - $265.11k
...rebuilding biotech for the AI era.When a breakthrough... ...and run AI agents and models directly in their... ...ll build the datasets, evaluations, and systems that help... ...distinguish strong model behavior.QUALIFICATIONS2+ years... ...program including equity, health, dental, vision, 401(k)...Work at officeLocal areaMonday to FridayShift work$60 - $100 per hour
...to improve how advanced AI models handle real-world accounting... ...correct solutions, and evaluation standards grounded in professional practice. Key... ...outputs, identifying missing behaviors, weak reasoning, misapplied... ..., reproductive health decisions, related medical...Hourly payFull timeLive inRelocationRelocation package$300k - $333k
...programs to align Gemini's style, tone, and behavior with the model specifications and taxonomy, scaling processes for consistent evaluations.Dive deep into technical details,... ...high tolerance for ambiguity inherent in AI research and operates at a fast pace. You...$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness... ...accumulation order, rounding behavior, and mixed-precision semantics across...Hourly payRemote work- ...Role Overview Help assess frontier AI systems at the boundary between legitimate... ...as benign, dual-use, or adversarial. Evaluate AI responses against a defined policy standard... ...should be able to distinguish routine professional requests from requests that may conceal...For contractorsRemote work
$60 - $80 per hour
...advanced large language models. In this role, you... ...campaign judgment to AI training data, partnering... ...Develop and improve evaluation guidelines and scoring... ...8 years of dedicated professional marketing experience,... ...childbirth, reproductive health decisions, or related...Hourly payWeekday work$100 per hour
...to help improve next-generation AI systems. This remote contract role focuses on evaluating psychological content,... ...ethical, and context-sensitive model training. Prior AI experience is... ...providing nuanced feedback and professional insights. Develop culturally...Hourly payContract workRemote work- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
$60 - $80 per hour
...building foundational AI models. You will create realistic marketing tasks, evaluate model outputs against... ...reasoning to improve model behavior and training data... ...8 years of dedicated professional experience in marketing... ...childbirth, reproductive health decisions, or related...Hourly payWeekday work$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous... ...judgment to the evaluation, design, and improvement... ...practice and ensure model outputs can withstand... ..., identifying missing behaviors, weak reasoning, unsupported... .... Hands-on professional use of large language...Hourly payFull timeLive inRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Behavioral Health Professional for AI Model Evaluation. Be the first to apply!





