AI Safety Practitioner - Expert Evaluator
Obsidian
We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. Responsibilities Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality. Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains. Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking. Identify unsafe outputs, hallucinations, reasoning failures, and policy violations. Provide structured feedback to improve model alignment and safety performance. Collaborate with AI researchers and safety teams on ongoing evaluation initiatives. Required Qualifications Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline. 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field. Excellent written English, critical thinking, and analytical reasoning skills. Ability to consistently evaluate nuanced and policy-sensitive scenarios. Preferred Qualifications Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation. Familiarity with safety policies, content moderation, or evaluation rubric development. Experience reviewing complex, high-risk, or ambiguous content. Why Join? Shape the safety and behaviour of frontier AI models used by millions worldwide. Work on challenging, real-world safety evaluations across nuanced and high-impact domains. Collaborate with leading AI researchers, engineers, and safety teams. #J-18808-Ljbffr Obsidian
- Handshake AI is seeking experienced CAD professionals with 2+ years of hands-on experience using SolidWorks and related CAD software, and professional proficiency in Mandarin Chinese, to support AI research through flexible, hourly contract work. This ongoing, project-...SuggestedRemote jobHourly payContract workFlexible hours
- A tech company focusing on AI research is looking for experienced Krita users for a flexible, project-based contract opportunity. This role allows you to earn while evaluating AI-generated content related to digital painting and concept art. Candidates should have at least...SuggestedRemote jobContract workFlexible hours
- ...Inpatient Nurses (RNs) to help train and evaluate AI systems used in clinical and healthcare... ...frontline experience to improve the accuracy, safety, and reliability of medical AI tools.... ...data for AI training datasets Provide expert feedback on nursing assessments and...Suggested
$30 per hour
About Prolific Prolific is not just another player in the AI space - we are building the biggest pool of quality human... ...Graphic and Visual Designers to act as Domain Experts for a high-level AI evaluation project. AI models are evolving beyond simple image generation...SuggestedRemote jobWork from homeFlexible hours- AuraOne is seeking an AI Policy Compliance AI Evaluator to perform remote adversarial evaluation, crafting 5-turn prompts and documenting failures with... ...a growing library of attack patterns to strengthen safety measures. The contractor role requires strong written communication...SuggestedRemote jobFor contractors
- AIUC is seeking experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models on policy-sensitive topics. You will assess AI-generated responses and provide structured feedback to improve model behavior. Join a collaborative team of...
- AuraOne is seeking a remote contractor to perform Chemical Safety Risk Evaluation by designing adversarial prompts, documenting failures with reproduction... ...per week. Applicants should have experience in red-teaming AI systems, familiarity with jailbreaking and policy-bypass...Remote jobFor contractors10 hours per week
- As part of AuraOne's safety-focused team, you will design 5‑turn attack scenarios, reproduce reported failures, and help harden AI systems before broader deployment. This is a contract position with US-eligibility suitable for remote workers. #J-18808-Ljbffr AuraOneRemote jobContract work
- TELUS Digital AI Community invites freelance content evaluators in the United States to join a remote independent contractor role. You will help improve AI-... ...handling sensitive material, and ensuring content meets guidelines to enhance safety and #J-18808-Ljbffr TELUS DigitalRemote jobFor contractorsFreelance
- Productive Playhouse seeks AI Evaluators to support evaluating AI chatbots by interacting with models, assessing capabilities, safety, and usefulness. This is a project-based, task-based engagement with flexible hours and batch deliveries. Open to freelancers outside the...Remote jobFreelanceFlexible hours
- Productive Playhouse is seeking Marathi-speaking AI Evaluators to test and assess next‑gen AI chatbots. You will interact with models, provide structured feedback on performance, safety, and usefulness, and deliver write-ups, ratings, and media per task. Remote, flexible...Remote jobFreelanceFlexible hours
- ...hiring PhD‑level biologists to help make advanced AI models safer. You'll apply your scientific expertise to evaluate and strengthen how these models handle... ...train you on the workflow. Responsibilities Write expert‑level prompts across specialized life‑science topics...Part timeImmediate start
- Mercor is seeking a remote Senior Red Team AI Specialist to test conversational agents and AI models against adversarial inputs. You will annotate vulnerabilities, surface systemic risks, and produce reproducible attack cases to help customers strengthen their AI systems...Remote work
- Visa Hunt seeks a Chemical Safety & Toxicology Expert (contractor, remote) to contribute to a client project enhancing chemical safety evaluation frameworks for AI training. You will apply domain knowledge in toxicology and regulated materials to shape model learning,...Remote jobFor contractors
- Mercor is building a remote red team to probe AI models with adversarial inputs. We focus on testing for jailbreaks, bias and safety vulnerabilities in conversational systems. You will generate actionable data and reports to help customers harden their AI. Ideal candidates...Remote job
- Mercor is building a red team for adversarial AI testing, reviewing outputs on sensitive topics with optional involvement in high-sensitivity projects, guided by clear guidelines and wellness resources. You will red team conversational AI models, generate high-quality...Remote job
- ...known as ActiveFence) is a leading trust, safety, and security company. Just like 'Alice'... ...the rabbit hole into the emerging world of AI and focus on safeguarding these... ...vulnerabilities. Vulnerability Assessment: Evaluate the security posture of AI models and infrastructure...Freelance
- AuraOne is seeking an Educational Safety AI Evaluator for a remote, contractor-based role. You will review educational safety AI outputs, apply the quality rubric, and produce auditable labels, rationales, and regression cases for AuraOne Human Data. Responsibilities include...Remote jobFor contractors
- ...is seeking experienced Adult Inpatient Nurses (RNs) to train and evaluate AI systems used in clinical and healthcare settings. This role leverages frontline nursing expertise to improve AI accuracy, safety, and reliability. You’ll work on projects requiring deep...
$80 - $120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ..., and Jack Dorsey . Position: Software / AI / IT / data Evaluator Type: Contract Compensation: $80–$120/hour Location...Contract workSummer workWork at officeRemote work- ...is seeking a General Nurse Subject Matter Expert for a remote contractor role to review... ...healthcare content. Your expertise will enhance AI models' accuracy in clinical reasoning.... ...model integrity through rigorous evaluation of responses. #J-18808-Ljbffr SME CareersRemote jobFor contractors
- ...the surgical setting. You will assess and implement anesthesia plans, monitor patients, and collaborate with physicians to ensure safety and quality. The role emphasizes compliance with HIPAA, regulatory standards, and continuous professional development. Prior CRNA experience...Part time
$280.79k - $312.58k
NYU Langone Health in New York seeks a Certified Registered Nurse Anesthetist to join the team. This role involves administering anesthesia, preparing for case management, and providing both pre and post-anesthetic patient care. Candidates must hold a Master's Degree from...$55 - $65 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Remote Role Responsibilities Review and evaluate AI-generated clinical outputs based on... ...for AI training datasets . Provide expert feedback on nursing assessments and documentation...Contract workSummer workRemote work- Obsidian is seeking a Spanish (Mexico) Audio Generalist Evaluator Expert for a high-impact audio AI research project. The role focuses on transcription, annotation, and evaluation tasks to help train advanced language models. Candidates need to be fluent in Spanish (Mexico...Temporary work10 hours per week
- About the role We are hiring expert Evaluators in Real estate / hospitality / events to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade...Hourly payWork at officeRemote work
$400 per month
Obsidian is seeking contributors for a Frontier Code Agents project, focused on evaluating AI coding models in fraud and risk engineering. Candidates will use AI coding tools to handle complex tasks and provide technical assessments. The role requires 2+ years of experience...$20 - $22 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Assamese Type: Contract Compensation: $2...Contract workSummer workRemote work$17 - $25 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Vietnamese Type: Contract Compensation:...Contract workSummer workRemote work$65 - $70 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Jack Dorsey . Position: Biology Expert (PhD) — AI Safety Type: Contract Compensation:... ...topics to enhance model understanding. Evaluate and annotate model responses for...Contract workSummer workImmediate startRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Safety Practitioner - Expert Evaluator. Be the first to apply!
- infection control practitioner New York, NY
- assistant practitioner New York, NY
- practitioner New York, NY
- holistic practitioner New York, NY
- subject matter expert New York, NY
- guest service support expert New York, NY
- technology expert New York, NY
- fulfillment expert New York, NY
- sql expert New York, NY
- program evaluator New York, NY
