Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Practitioner - Expert Evaluator

Obsidian

We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. Responsibilities Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality. Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains. Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking. Identify unsafe outputs, hallucinations, reasoning failures, and policy violations. Provide structured feedback to improve model alignment and safety performance. Collaborate with AI researchers and safety teams on ongoing evaluation initiatives. Required Qualifications Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline. 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field. Excellent written English, critical thinking, and analytical reasoning skills. Ability to consistently evaluate nuanced and policy-sensitive scenarios. Preferred Qualifications Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation. Familiarity with safety policies, content moderation, or evaluation rubric development. Experience reviewing complex, high-risk, or ambiguous content. Why Join? Shape the safety and behaviour of frontier AI models used by millions worldwide. Work on challenging, real-world safety evaluations across nuanced and high-impact domains. Collaborate with leading AI researchers, engineers, and safety teams. #J-18808-Ljbffr Obsidian

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Safety Practitioner - Expert Evaluator in New York, NY vacancy
  • Handshake AI is seeking experienced CAD professionals with 2+ years of hands-on experience using SolidWorks and related CAD software, and professional proficiency in Mandarin Chinese, to support AI research through flexible, hourly contract work. This ongoing, project-... 
    Suggested
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Handshake AI

    New York, NY
    2 days ago
  • A tech company focusing on AI research is looking for experienced Krita users for a flexible, project-based contract opportunity. This role allows you to earn while evaluating AI-generated content related to digital painting and concept art. Candidates should have at least... 
    Suggested
    Remote job
    Contract work
    Flexible hours

    Handshake

    New York, NY
    18 hours ago
  •  ...Inpatient Nurses (RNs) to help train and evaluate AI systems used in clinical and healthcare...  ...frontline experience to improve the accuracy, safety, and reliability of medical AI tools....  ...data for AI training datasets Provide expert feedback on nursing assessments and... 
    Suggested

    Mercor

    New York, NY
    4 days ago
  • $30 per hour

    About Prolific Prolific is not just another player in the AI space - we are building the biggest pool of quality human...  ...Graphic and Visual Designers to act as Domain Experts for a high-level AI evaluation project. AI models are evolving beyond simple image generation... 
    Suggested
    Remote job
    Work from home
    Flexible hours

    Prolific

    New York, NY
    2 days ago
  • AuraOne is seeking an AI Policy Compliance AI Evaluator to perform remote adversarial evaluation, crafting 5-turn prompts and documenting failures with...  ...a growing library of attack patterns to strengthen safety measures. The contractor role requires strong written communication... 
    Suggested
    Remote job
    For contractors

    AuraOne

    New York, NY
    1 day ago
  • AIUC is seeking experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models on policy-sensitive topics. You will assess AI-generated responses and provide structured feedback to improve model behavior. Join a collaborative team of... 

    Dorado

    New York, NY
    18 hours ago
  • AuraOne is seeking a remote contractor to perform Chemical Safety Risk Evaluation by designing adversarial prompts, documenting failures with reproduction...  ...per week. Applicants should have experience in red-teaming AI systems, familiarity with jailbreaking and policy-bypass... 
    Remote job
    For contractors
    10 hours per week

    AuraOne

    New York, NY
    3 days ago
  • As part of AuraOne's safety-focused team, you will design 5‑turn attack scenarios, reproduce reported failures, and help harden AI systems before broader deployment. This is a contract position with US-eligibility suitable for remote workers. #J-18808-Ljbffr AuraOne
    Remote job
    Contract work

    AuraOne

    New York, NY
    1 day ago
  • TELUS Digital AI Community invites freelance content evaluators in the United States to join a remote independent contractor role. You will help improve AI-...  ...handling sensitive material, and ensuring content meets guidelines to enhance safety and #J-18808-Ljbffr TELUS Digital
    Remote job
    For contractors
    Freelance

    TELUS Digital

    New York, NY
    1 day ago
  • Productive Playhouse seeks AI Evaluators to support evaluating AI chatbots by interacting with models, assessing capabilities, safety, and usefulness. This is a project-based, task-based engagement with flexible hours and batch deliveries. Open to freelancers outside the... 
    Remote job
    Freelance
    Flexible hours

    Triwill Group

    New York, NY
    2 days ago
  • Productive Playhouse is seeking Marathi-speaking AI Evaluators to test and assess next‑gen AI chatbots. You will interact with models, provide structured feedback on performance, safety, and usefulness, and deliver write-ups, ratings, and media per task. Remote, flexible... 
    Remote job
    Freelance
    Flexible hours

    Triwill Group

    New York, NY
    2 days ago
  •  ...hiring PhD‑level biologists to help make advanced AI models safer. You'll apply your scientific expertise to evaluate and strengthen how these models handle...  ...train you on the workflow. Responsibilities Write expert‑level prompts across specialized life‑science topics... 
    Part time
    Immediate start

    Obsidian

    New York, NY
    4 days ago
  • Mercor is seeking a remote Senior Red Team AI Specialist to test conversational agents and AI models against adversarial inputs. You will annotate vulnerabilities, surface systemic risks, and produce reproducible attack cases to help customers strengthen their AI systems... 
    Remote work

    Neon

    New York, NY
    2 days ago
  • Visa Hunt seeks a Chemical Safety & Toxicology Expert (contractor, remote) to contribute to a client project enhancing chemical safety evaluation frameworks for AI training. You will apply domain knowledge in toxicology and regulated materials to shape model learning,... 
    Remote job
    For contractors

    Visa Hunt

    New York, NY
    2 days ago
  • Mercor is building a remote red team to probe AI models with adversarial inputs. We focus on testing for jailbreaks, bias and safety vulnerabilities in conversational systems. You will generate actionable data and reports to help customers harden their AI. Ideal candidates... 
    Remote job

    Mercor

    New York, NY
    3 days ago
  • Mercor is building a red team for adversarial AI testing, reviewing outputs on sensitive topics with optional involvement in high-sensitivity projects, guided by clear guidelines and wellness resources. You will red team conversational AI models, generate high-quality... 
    Remote job

    Neon

    New York, NY
    2 days ago
  •  ...known as ActiveFence) is a leading trust, safety, and security company. Just like 'Alice'...  ...the rabbit hole into the emerging world of AI and focus on safeguarding these...  ...vulnerabilities. Vulnerability Assessment: Evaluate the security posture of AI models and infrastructure... 
    Freelance

    Socket.dev

    New York, NY
    2 days ago
  • AuraOne is seeking an Educational Safety AI Evaluator for a remote, contractor-based role. You will review educational safety AI outputs, apply the quality rubric, and produce auditable labels, rationales, and regression cases for AuraOne Human Data. Responsibilities include... 
    Remote job
    For contractors

    AuraOne

    New York, NY
    1 day ago
  •  ...is seeking experienced Adult Inpatient Nurses (RNs) to train and evaluate AI systems used in clinical and healthcare settings. This role leverages frontline nursing expertise to improve AI accuracy, safety, and reliability. You’ll work on projects requiring deep... 

    Mercor

    New York, NY
    4 days ago
  • $80 - $120 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ..., and Jack Dorsey . Position: Software / AI / IT / data Evaluator Type: Contract Compensation: $80–$120/hour Location... 
    Contract work
    Summer work
    Work at office
    Remote work

    Mercor

    New York, NY
    4 days ago
  •  ...is seeking a General Nurse Subject Matter Expert for a remote contractor role to review...  ...healthcare content. Your expertise will enhance AI models' accuracy in clinical reasoning....  ...model integrity through rigorous evaluation of responses. #J-18808-Ljbffr SME Careers
    Remote job
    For contractors

    SME Careers

    New York, NY
    2 days ago
  •  ...the surgical setting. You will assess and implement anesthesia plans, monitor patients, and collaborate with physicians to ensure safety and quality. The role emphasizes compliance with HIPAA, regulatory standards, and continuous professional development. Prior CRNA experience... 
    Part time

    Highmark Health

    New York, NY
    8 hours ago
  • $280.79k - $312.58k

    NYU Langone Health in New York seeks a Certified Registered Nurse Anesthetist to join the team. This role involves administering anesthesia, preparing for case management, and providing both pre and post-anesthetic patient care. Candidates must hold a Master's Degree from...

    NYU Langone Health

    New York, NY
    1 day ago
  • $55 - $65 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Remote Role Responsibilities Review and evaluate AI-generated clinical outputs based on...  ...for AI training datasets . Provide expert feedback on nursing assessments and documentation... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    6 days ago
  • Obsidian is seeking a Spanish (Mexico) Audio Generalist Evaluator Expert for a high-impact audio AI research project. The role focuses on transcription, annotation, and evaluation tasks to help train advanced language models. Candidates need to be fluent in Spanish (Mexico... 
    Temporary work
    10 hours per week

    Obsidian

    New York, NY
    4 days ago
  • About the role We are hiring expert Evaluators in Real estate / hospitality / events to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade... 
    Hourly pay
    Work at office
    Remote work

    Obsidian

    New York, NY
    3 days ago
  • $400 per month

    Obsidian is seeking contributors for a Frontier Code Agents project, focused on evaluating AI coding models in fraud and risk engineering. Candidates will use AI coding tools to handle complex tasks and provide technical assessments. The role requires 2+ years of experience... 

    Obsidian

    New York, NY
    18 hours ago
  • $20 - $22 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Assamese Type: Contract Compensation: $2... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    1 day ago
  • $17 - $25 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Vietnamese Type: Contract Compensation:... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    6 days ago
  • $65 - $70 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Jack Dorsey . Position: Biology Expert (PhD) — AI Safety Type: Contract Compensation:...  ...topics to enhance model understanding. Evaluate and annotate model responses for... 
    Contract work
    Summer work
    Immediate start
    Remote work

    Mercor

    New York, NY
    23 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Practitioner - Expert Evaluator. Be the first to apply!