AI Safety Policy Evaluator, Violence & Threats
$93.6k - $114.4kJobgether
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Safety Policy Evaluator, Violence & Threats based in United States.
This role sits at the intersection of AI safety, content policy, and expert human judgment, helping improve how advanced AI models handle violent and threatening content.
You’ll evaluate user requests, model responses, and conversation context to distinguish legitimate fictional, educational, historical, or defensive content from material that could enable real-world harm.
The work focuses heavily on nuanced edge cases where intent, context, and a single detail can materially change the appropriate policy decision.
You’ll contribute not only to evaluations, but also to adversarial testing, policy refinement, calibration, and the identification of emerging safety gaps.
This is a non-engineering role for someone with deep experience in areas such as violent fiction, military or emergency response, crisis intervention, threat assessment, trust and safety, or related fields.
You’ll work in a structured, feedback-rich environment where clear reasoning, consistency, and the ability to separate personal beliefs from policy standards are essential.
Because the role involves regular exposure to difficult material, candidates must be prepared to engage with sensitive content carefully, professionally, and sustainably.
- Evaluate user requests, AI responses, and conversation histories involving violence, weapons, threats, self-harm, dark fiction, and related sensitive topics.
- Distinguish fictional, educational, historical, journalistic, defensive, or expressive content from requests that meaningfully facilitate real-world harm.
- Assess whether an AI response provides actionable real-world capability, regardless of how the original request is framed.
- Differentiate ordinary anger, frustration, venting, or dark humor from credible threats and potential crisis indicators.
- Apply relevant customer policies consistently while considering intent, context, precedent, and established team guidance rather than relying on rote annotation.
- Select defensible classifications for ambiguous cases and produce concise, well-supported rationales referencing policy language and relevant conversation details.
- Develop and refine adversarial or borderline prompts that test how AI systems handle difficult policy boundaries.
- Identify policy gaps, contradictions, recurring ambiguities, and emerging edge cases, escalating findings to project leads and policy teams.
- Participate actively in calibration and adjudication discussions, respectfully challenging interpretations and updating judgments when stronger reasoning emerges.
- Maintain high accuracy, consistency, and attention to detail across repetitive, feedback-heavy evaluation workflows.
- Contribute to improving AI safety standards by turning nuanced human judgment into clear, auditable evaluation guidance.
Requirements
- Demonstrated depth of experience in at least one relevant domain, such as violent fiction, game design or game mastering, film and television, military or law enforcement, emergency medicine, crisis counseling, threat assessment, trust and safety, content moderation, journalism, or law.
- Strong ability to distinguish fictional or contextualized depictions of violence from content that facilitates or signals real-world harm.
- Demonstrated understanding of how intent, context, language, and subtle changes in a request can materially affect a safety assessment.
- Ability to make nuanced judgment calls while separating personal beliefs from the policy standard being applied.
- Strong written communication skills, with the ability to explain complex decisions clearly enough for another evaluator to audit the reasoning.
- Ability to remain open-minded, challenge assumptions constructively, and revise conclusions when stronger evidence or reasoning emerges.
- Strong attention to detail and consistency when working through repeated evaluations involving difficult or sensitive material.
- Familiarity with AI tools and a strong interest in understanding where language models may over-refuse, under-refuse, or misunderstand user intent.
- Prior experience with AI evaluation, red teaming, data annotation, RLHF, trust and safety, content moderation, or language-model assessment is helpful but not required.
- Experience with calibration sessions, inter-rater agreement, adjudication workflows, or structured policy evaluation is a plus.
- Additional valuable experience may include published or produced work involving violence, military or security experience, crisis intervention, threat assessment, forensic or clinical psychology, violence prevention, weapons disciplines, or professional experience evaluating ChatGPT, Claude, Gemini, or similar AI systems.
- A degree, security clearance, or technical/software engineering background is not required.
- Ability to work remotely in the United States on a Monday-Friday schedule from 8:00 AM to 5:00 PM PT.
Benefits
- Compensation: $45–$55 per hour.
- Employment classification: W-2.
- Work arrangement: Fully remote within the United States.
- Schedule: Monday through Friday, 8:00 AM–5:00 PM Pacific Time.
- Assignment: Ongoing opportunity with a planned start date of September 21, 2026.
- Benefits eligibility: Eligible for available employee benefits.
- Structured evaluation frameworks, professional guidelines, content rotation, and exposure limits designed to support sustainable work with sensitive material.
- Access to mental health support given the nature of the content reviewed.
- Opportunity to contribute directly to the safety, reliability, and responsible development of advanced AI systems.
- Exposure to complex policy questions and emerging challenges at the intersection of AI, violence, threats, and content safety.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
$45 - $55 per hour
...non-engineering content-policy evaluation role. Applicants must demonstrate... ...response, crisis or threat assessment, trust and safety, content moderation, or... ...25, we started Handshake AI and built the fastest-growing... ...Policy Specialist on the Violence & Fiction team, you will...PolicyRemote jobMonday to FridayShift work- Handshake is seeking an AI Policy Specialist on the Violence & Fiction team to evaluate user requests and model responses within a full conversation context, distinguishing violence in fiction from real-world uplift and ensuring precise policy application. You will read...PolicyRemote job
- ...Threat Intelligence AI Evaluator is a remote evaluation track for reviewing threat intelligence ai evaluation... ...an unsafe response with the correct policy category and severity. Audit a 50-... ..., content moderation, or trust & safety review. Experience with inter-rater...PolicyRemote jobHourly payFor contractors10 hours per week
$43 - $47 per hour
...Overview Help improve how advanced AI models respond to sensitive topics in... ...language fluency and cultural judgment to evaluate model behavior and support safer,... ...content. Experience in trust and safety, content moderation, policy evaluation, or adversarial testing....PolicyRemote jobHourly payImmediate start$18 - $22 per hour
...Role Overview Help improve the safety of advanced AI models by applying Thai language fluency and cultural judgment to evaluate how models respond to sensitive topics in Thai. Training... ...in trust and safety, content moderation, policy evaluation, or adversarial testing....PolicyRemote jobHourly payImmediate start- ...AI Safety and Red-Team Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios... ...prompts that probe known weakness classes (jailbreak, policy bypass, prompt injection) for AI Safety and Red-Team...PolicyRemote jobHourly payFor contractors10 hours per week
$43 - $47 per hour
...Role Overview Help evaluate and strengthen how advanced AI systems respond to sensitive topics in Spanish. This role combines Spanish... ...or technical content. Background in trust and safety, content moderation, policy evaluation, or adversarial testing. Work Terms...PolicyRemote jobHourly payImmediate start$48 - $52 per hour
...Role Overview Help improve the safety of advanced AI systems by applying Chinese language fluency and cultural judgment to evaluate how models respond to sensitive topics. Training... ...in trust and safety, content moderation, policy evaluation, or adversarial testing....PolicyRemote jobHourly payImmediate start- ...Healthcare Compliance AI Evaluator is a remote clinical-review track for evaluating AI outputs... ..., and guideline adherence; flag patient-safety issues; and document the corrected... ...Clinical review Legal reasoning Policy review Risk analysis Healthcare...PolicyRemote jobHourly payFor contractors10 hours per week
- OpenAI is seeking a Technical Threat Investigator to protect our systems by conducting deep investigations into sophisticated threat... ...findings into durable tooling, and collaborate with Safety, Product Policy, and Integrity teams to disrupt adversarial activity. Responsibilities...PolicyRemote job
- Handshake is seeking an AI Image Evaluator to assess prompts and generated images for quality, accuracy, and policy compliance. You will weigh visual details, compare outputs, and explain your reasoning to help train evaluation models. Remote US role with flexible scheduling...PolicyRemote jobFlexible hours
- OpenTrain AI is seeking an insurance policy operations specialist to craft high-quality reasoning data for AI evaluation. You will design realistic workflows across the policy lifecycle, produce reference answers, and assess model outputs against operational standards....PolicyFor contractorsRemote workFlexible hours
$130k - $160k
...System Security Officer (ISSO) / Control Evaluator - High/Senior to support a... ...frameworks, federal information security policy, security assessment methodologies, and... ...protection (both internal and external threats), and collaborate on building out risk register...PolicyFull timeWork experience placementWork at officeLocal areaRemote workFlexible hours$84.18k - $142.37k
...Functional Title: Forensic Evaluator - State Hospital - Maximum Security... ...evaluations, as well as violence risk assessments under 46C. Responsibilities... ...accordance with agency leave policy and performs other duties as... ...supports hospital and agency safety (including patient safety),...PolicyFull timeTemporary workPart timeTraineeshipInternshipWork at officeRemote workShift workRotating shift- ...Patent Claim AI Evaluator is a remote review track for evaluating AI outputs in patent law workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can train on it. Why...PolicyRemote jobHourly payFor contractors10 hours per week
- ...Assessment Rubric AI Evaluator is a remote review track for evaluating AI outputs across assessment rubric ai specialist operations workflows. Reviewers grade workflow correctness, policy adherence, and stakeholder fit; flag operational risk; and document the right next...PolicyRemote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Procurement Compliance AI Evaluator is a remote review track for evaluating AI outputs in compliance workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can train on...PolicyRemote jobHourly payFor contractors10 hours per week
- ...Harmful Instruction Refusal Risk Evaluator is a remote red-team track for stress-testing AI systems against adversarial... ...rubric clause it violated so the safety team can patch the gap. Why this... ...known weakness classes (jailbreak, policy bypass, prompt injection) for...PolicyRemote jobHourly payFor contractors10 hours per week
$80 - $120 per hour
...Compliance / regulatory response with financial-services AI Evaluator is a remote review track for evaluating AI outputs across finance... ...workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and...PolicyRemote jobFor contractorsWork experience placement10 hours per week$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours- ...team ensures the physical safety and security of the... ...Protective Intelligence & Threat Analyst, you will... ...threat investigations to evaluate the credibility, intent... ...Python, data platforms, AI-enabled tools, and other... ...Employment Opportunity Policy Statement. Background...PolicyWork at officeRemote workRelocationNight shift
$100k - $110k
...Data Evaluator/Analyst MELE Associates, Inc. is seeking to add an experienced Data Evaluator/Analyst to support the Department of Energy... ...: Bachelor's degree in economics, statistics, public policy, data science, or a related field (or equivalent combination of...PolicyContract workFor contractorsWork at officeRemote work$8.18k - $11.86k
...Functional Title: Psychologist III Forensic Evaluator Job Title: Psychologist III Agency:... ...under Texas Family Code Chapter 55, and violence risk assessments under CCP 46C across... ...with applicable statutes, standards, and policies. Evaluates and ensures the professional...PolicyFull timeTemporary workPart timeInternshipWork at officeRemote workShift work- ...The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent... ...confidentiality. Adhere to hospital and departmental policies and state and federal regulations. Orient new staff and students...Policy
- ...Epidemiologist who will work on program evaluation and data-related activities for multiple... ...Resilience programs, Implementation of SUD in Violence Intervention and Prevention (VIP)... ...specific analyses that inform program and policy development, with a focus on racial and...PolicyFull timeWork experience placementWork at officeFlexible hours
- Moonshot is seeking a Head of AI Safety to lead the delivery,... ...portfolio, combining expertise in violence prevention, behavioral risk,... ...senior role collaborates with policy, research, engineering and product... ...teams to steer red teaming, evaluation frameworks, and client...Policy
$16 - $20 per hour
...of Nursing is hiring part-time research evaluators for projects focused on dementia care in... ...check(s) in accordance with University policies. CAMPUS SECURITY CRIME STATISTICS... ...combined Annual Security and Annual Fire Safety Report (ASR). The ASR includes crime statistics...PolicyHourly payPart timeFor contractorsSummer workRemote work- ...MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers...Contract workTemporary workImmediate startRemote work
- ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial reasoning and quantitative analysis through flexible, contract-based work. Key responsibilities Assessing the factuality and...Full timeContract workRemote workFlexible hours
$14.5 per hour
A technology company is seeking an AI Web Search Evaluator to enhance the quality of search engine results. This flexible, remote, part-time role focuses on analyzing search performance and providing feedback to improve algorithms. Ideal candidates will have strong analytical...Hourly payPart timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Safety Policy Evaluator, Violence & Threats. Be the first to apply!




