Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Practitioner - Expert Evaluator

Obsidian

We are seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across complex, policy-sensitive, and ambiguous ("grey-area") topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations and feedback. Responsibilities Evaluate AI-generated responses for safety, factual accuracy, policy compliance, and overall quality. Review content involving misinformation, political persuasion, self-harm, violence, cyber, biosecurity, and other sensitive domains. Apply and refine evaluation rubrics for RLHF, SFT, and AI safety benchmarking. Identify unsafe outputs, hallucinations, reasoning failures, and policy violations. Provide structured feedback to improve model alignment and safety performance. Collaborate with AI researchers and safety teams on ongoing evaluation initiatives. Required Qualifications Bachelor's degree or higher in Journalism, Communications, Psychology, Sociology, Public Policy, Law, Biology, Chemistry, Computer Science, or a related discipline. 5+ years of professional experience in AI Safety, Trust & Safety, journalism, public policy, scientific research, security, or a related field. Excellent written English, critical thinking, and analytical reasoning skills. Ability to consistently evaluate nuanced and policy-sensitive scenarios. Preferred Qualifications Experience with AI Safety, RLHF, SFT, Trust & Safety, or AI evaluation. Familiarity with safety policies, content moderation, or evaluation rubric development. Experience reviewing complex, high-risk, or ambiguous content. Why Join? Shape the safety and behaviour of frontier AI models used by millions worldwide. Work on challenging, real-world safety evaluations across nuanced and high-impact domains. Collaborate with leading AI researchers, engineers, and safety teams. #J-18808-Ljbffr Obsidian

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Safety Practitioner - Expert Evaluator in New York, NY vacancy
  • Mercor is seeking experienced musicians to evaluate generative music AI models, collaborating with a leading AI lab. You will assess AI-generated music and lyrics in Telugu and English, applying detailed quality standards across genres. Responsibilities include comparing... 
    Suggested
    Part time
    Immediate start

    Obsidian

    New York, NY
    2 days ago
  • Mercor is seeking experienced Medical and Health Services Managers to evaluate and improve AI-generated healthcare operations content and workflows. You will leverage your expertise in directing clinical services, personnel, budgets, and compliance across healthcare facilities... 
    Suggested

    Mercor

    New York, NY
    3 days ago
  • Mercor is hiring Legal Experts to evaluate AI-generated responses for employment and labor law scenarios. This fully remote, hourly contract offers flexible 6-15 hours per week. You will assess accuracy, provide feedback to improve model behavior and participate in calibration... 
    Suggested
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Intellerzone

    New York, NY
    3 days ago
  • $30 per hour

     ...About Prolific Prolific is not just another player in the AI space – we are building the biggest pool of quality human...  ...Graphic and Visual Designers to act as Domain Experts for a high-level AI evaluation project. AI models are evolving beyond simple image generation... 
    Suggested
    Remote work
    Work from home
    Flexible hours

    Prolific

    New York, NY
    3 hours ago
  • Mercor is seeking experienced musicians to evaluate generative music AI models in collaboration with a leading AI lab. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Punjabi and English. Requirements include... 
    Suggested

    Mercor

    New York, NY
    3 days ago
  •  ...Inpatient Nurses (RNs) to help train and evaluate AI systems used in clinical and healthcare...  ...frontline experience to improve the accuracy, safety, and reliability of medical AI tools....  ...data for AI training datasets Provide expert feedback on nursing assessments and... 

    Mercor

    New York, NY
    4 days ago
  • AIUC is seeking experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models on policy-sensitive topics. You will assess AI-generated responses and provide structured feedback to improve model behavior. Join a collaborative team of... 

    Dorado

    New York, NY
    23 hours ago
  •  ...Mercor, we believe the safest AI is the one that’s already been...  ...for this project - human data experts who probe AI models with adversarial...  ...customer AI systems Evaluation coverage expands: more scenarios...  ...production Mercor customers trust the safety of their AI because you’ve... 
    Remote work

    Mercor Inc

    New York, NY
    3 days ago
  • Mercor seeks experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models across policy-sensitive topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured evaluations... 

    Mercor

    New York, NY
    1 day ago
  • Visa Hunt is seeking a Child & Online Safety Expert to contribute to a high-impact project focused on online safety and youth mental-health...  ...contractor arrangement allows you to apply domain knowledge to train AI systems and shape how models learn. You will develop taxonomies... 
    Remote job
    For contractors

    Visa Hunt

    New York, NY
    2 days ago
  • We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area"... 

    Obsidian

    New York, NY
    4 days ago
  • Obsidian is seeking experienced AI Safety Practitioners to evaluate the safety, quality, and alignment of frontier AI models across policy-sensitive topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured... 

    Obsidian

    New York, NY
    2 days ago
  •  ...is seeking experienced Adult Inpatient Nurses (RNs) to train and evaluate AI systems used in clinical and healthcare settings. This role leverages frontline nursing expertise to improve AI accuracy, safety, and reliability. You’ll work on projects requiring deep... 

    Mercor

    New York, NY
    4 days ago
  •  ...is seeking a General Nurse Subject Matter Expert for a remote contractor role to review...  ...healthcare content. Your expertise will enhance AI models' accuracy in clinical reasoning....  ...model integrity through rigorous evaluation of responses. #J-18808-Ljbffr SME Careers
    Remote job
    For contractors

    SME Careers

    New York, NY
    2 days ago
  •  ...the surgical setting. You will assess and implement anesthesia plans, monitor patients, and collaborate with physicians to ensure safety and quality. The role emphasizes compliance with HIPAA, regulatory standards, and continuous professional development. Prior CRNA experience... 
    Part time

    Highmark Health

    New York, NY
    23 hours ago
  • $280.79k - $312.58k

    NYU Langone Health in New York seeks a Certified Registered Nurse Anesthetist to join the team. This role involves administering anesthesia, preparing for case management, and providing both pre and post-anesthetic patient care. Candidates must hold a Master's Degree from...

    NYU Langone Health

    New York, NY
    1 day ago
  •  ...About the role We are hiring expert Evaluators in Clinical / biomedical / pharma to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade... 
    Hourly pay
    Work at office
    Remote work

    Mercor Inc

    New York, NY
    1 day ago
  • $400 per month

    Obsidian is seeking contributors for a Frontier Code Agents project, focused on evaluating AI coding models in fraud and risk engineering. Candidates will use AI coding tools to handle complex tasks and provide technical assessments. The role requires 2+ years of experience... 

    Obsidian

    New York, NY
    23 hours ago
  • Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. Start date: immediate. Duration: up to... 
    Hourly pay
    Immediate start

    Obsidian

    New York, NY
    4 days ago
  • $100 - $150 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...reviewed judgment. Preferred Prior experience with AI training, evaluation, or human-data projects. #J-18808-Ljbffr United States Digital... 
    Summer work
    Remote work

    United States Digital Space LLC

    New York, NY
    4 days ago
  • YO AI Labs is seeking an established Physics Professor/PI to provide high-level scientific expertise for AI-system evaluation. Remote contractor role focusing on evaluating physics arguments, testing claims with rigorous judgments, and drafting defensible written analyses... 
    Remote job
    For contractors

    YO AI Labs

    New York, NY
    3 days ago
  • Join a growing network of industry experts supporting AI research and development. Overview We\'re building a select group of experienced healthcare professionals to help evaluate and improve how AI systems understand and reason about healthcare topics. You\'ll bring your... 
    Contract work
    Part time
    Remote work
    10 hours per week
    Flexible hours

    Obsidian

    New York, NY
    2 days ago
  •  ...MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers... 
    Contract work
    Temporary work
    Immediate start
    Remote work

    MERIT Beauty

    New York, NY
    23 hours ago
  •  ...Job Description Job Description Developer & Infrastructure Expert Role Type: Contractor Location: Remote Job Overview...  ...are seeking experienced Developer & Infrastructure Experts to evaluate AI-powered workflows across software development, cloud... 
    Remote job
    For contractors

    YO AI Labs

    New York, NY
    11 days ago
  •  ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional... 
    Work at office
    Remote work

    Obsidian

    New York, NY
    2 days ago
  •  ...BAM Ventures is seeking Swedish-speaking remote annotators to evaluate AI-generated content, ensuring that it's coherent and aligns with real-world expectations. Your role will involve reviewing outputs, identifying deviations, and providing structured feedback to enhance... 
    Remote work

    BAM VENTURES LLC

    New York, NY
    2 days ago
  • 1. Role Overview Mercor is partnering with a leading AI research organization to engage experienced sales engineers for a project focused on evaluating how well AI systems perform real-world technical sales work. Rather than producing deliverables yourself, you will define... 

    Obsidian

    New York, NY
    4 days ago
  • Obsidian is seeking experienced evaluators to test AI systems on complex personal workflows across health, travel, planning, home services, and career search. You will create realistic prompts, execute tasks with screen recording, and use your own plugins to complete actions... 

    Obsidian

    New York, NY
    1 day ago
  • Mercor seeks experienced brand designers to define grading criteria for real-world brand deliverables and to score AI-generated versus human work with detailed written justifications. You will apply evidence-based judgment, ensure reproducibility, and iteratively incorporate... 

    Obsidian

    New York, NY
    2 days ago
  • $20 - $30 per hour

    A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong... 
    Remote job
    Hourly pay

    Crossing Hurdles

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Practitioner - Expert Evaluator. Be the first to apply!