Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Trust and Safety Policy AI Evaluator

$65 - $70 per hour

AI Trainer Jobs

Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap. Category: AI Safety & Red Teaming · Pay: $65–$70 / hr · Location: Remote — US-eligible · Contractor Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. About the role Trust and Safety Policy AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap. Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them. Push on refusal boundaries and dual-use risk before a model ships. Responsibilities Design adversarial prompts that probe known weakness classes (jailbreak, policy bypass, prompt injection) for Trust and Safety Policy AI Evaluator assignments. Document every successful attack with reproduction steps and the policy clause it violated. Score model defenses across single-turn and multi-turn conversations. Triage emerging attack vectors and route them to the safety team with severity ratings. Maintain a personal library of attack patterns and propose new red-team rubrics. Role details Track Adversarial evaluation Work model Remote · Independent specialist contractor Compensation Hourly rate confirmed after the interview process. Eligible from US What you should bring Demonstrated experience red-teaming AI systems, security research, or adversarial ML work for Trust and Safety Policy AI Evaluator work. Strong written communication — your reports become the patch ticket. Comfort working in policy-grey areas with clear documentation of what was attempted and why. Familiarity with prompt-injection, jailbreak, and policy-bypass taxonomies. Reliable async availability for at least 10 hours per week. Role signals Example tasks Construct a 5-turn adversarial conversation that bypasses a specific policy clause and write up the patch ticket. Score a model's defenses against a known jailbreak pattern across 20 variants. Propose a new red-team rubric category after spotting an emerging attack vector. Reproduce a failure another reviewer reported and confirm the severity tag. Useful experience Background in offensive security, AppSec, or trust & safety operations. Experience publishing or reproducing public adversarial-ML research. Multilingual fluency for cross-language attack testing. Compensation and schedule Hourly rate confirmed after the interview process. Expected arrangement: contractor , with program-defined task volume and review pacing. Placement depends on current program demand and reviewer confirmation. Skills used in matching Adversarial prompting Red-team analysis Policy taxonomy Failure documentation Trust and Safety Policy AI evaluation Legal reasoning Policy review Risk analysis Trust Safety Policy #J-18808-Ljbffr AI Trainer Jobs

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Trust and Safety Policy AI Evaluator in New York, NY vacancy
  • AuraOne is seeking a Chemicals Safety Risk Evaluator in a remote, US-eligible red-team track to stress-test AI systems against adversarial prompts. Reviewers craft attack...  ...tie each successful jailbreak to the violated policy clause so the safety team can patch gaps. As... 
    Policy
    Remote job
    For contractors

    AI Trainer Jobs

    New York, NY
    4 days ago
  • About the role Hindi Localization AI Evaluator is a remote evaluation track for reviewing hindi...  ...an unsafe response with the correct policy category and severity. Audit a 50-row...  ...in linguistics, content moderation, or trust & safety review. Experience with inter-rater agreement... 
    Policy
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    4 days ago
  • About the role Swahili Evaluation AI Evaluator is a remote evaluation track for reviewing swahili...  ...an unsafe response with the correct policy category and severity. Audit a 50-row...  ...in linguistics, content moderation, or trust & safety review. Experience with inter-rater agreement... 
    Policy
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    4 days ago
  • TikTok's Trust & Safety team works to create a safe and trusted platform for billions of people around the world. We combine policy, technology, and operations to detect harmful content, protect...  ...reporting and observability.- Explore AI-powered solutions for monitoring... 
    Policy

    TikTok

    New York, NY
    13 hours ago
  • AI Trainer Jobs is seeking a Grocery Shopper Expert contractor for remote work in the US. You will review AI outputs across grocery shopper operations, grade workflow correctness, policy adherence, and stakeholder fit, and document next steps for model training. Experience... 
    Policy
    Remote job
    For contractors
    10 hours per week

    AI Trainer Jobs

    New York, NY
    3 hours ago
  • Torsap Thai Kitchen seeks an experienced Director of Engineering, Trust & Safety to lead the technical strategy and platforms safeguarding its...  ...and executive communication. You will collaborate with Legal, Policy, and Product teams to translate risks into robust solutions... 
    Policy

    Torsap Thai Kitchen

    New York, NY
    3 hours ago
  • AuraOne is seeking a Children Safety Content AI Evaluator for a remote, contractor-based role. You will evaluate model outputs, compare responses, and label content with structured tags according to our quality rubric. Ideal candidates have prior evaluation/annotation experience... 
    Remote job
    For contractors
    10 hours per week

    AI Trainer Jobs

    New York, NY
    4 days ago
  • AuraOne is seeking a Manipulation Risk Risk Evaluator for a remote, contractor role focused on stress-testing AI systems against adversarial prompts. Reviewers craft...  ...jailbreak with the violated rubric clause so the safety team can patch gaps. Adversarial evaluation... 
    For contractors
    Remote work

    AI Trainer Jobs

    New York, NY
    4 days ago
  • $70k - $90k

     ...values, and foster a culture of inclusion, safety, and growth. As an organization, we...  ...employees and supervisors to observe and evaluate their knowledge, skills, and abilities to...  ...to the Company business and operational policies and procedures. Assist in reviewing, improving... 
    Policy
    Full time
    Temporary work
    For contractors
    Local area
    Relocation

    18 Primoris Renewable Energy

    New York, NY
    1 day ago
  • $75 - $110 per hour

    AI Trainer Jobs seeks a Nursing Triage Safety Evaluator to remotely review AI outputs that touch nursing. You will assess differential reasoning, dosing logic, and guideline adherence, flag patient-safety issues, and document corrected clinical reasoning so the modeling... 
    Remote job
    Hourly pay
    For contractors

    AI Trainer Jobs

    New York, NY
    4 days ago
  • $30 - $60 per hour

    Web Browsing Evaluator is a remote evaluation track for reviewing web...  ...team can use to retrain. AI data reviewers help turn web...  ...unsafe response with the correct policy category and severity. Audit...  ...linguistics, content moderation, or trust & safety review. Experience with inter... 
    Policy
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    4 days ago
  • AI Trainer Jobs is seeking an Equity Research AI Evaluator to remotely review AI outputs in finance and risk workflows. You will grade calculations, narrative reasoning, and policy adherence, flag issues, and document corrected workpapers for training. Prerequisites include... 
    Policy
    Remote job
    For contractors

    AI Trainer Jobs

    New York, NY
    1 day ago
  •  ...Invoice Understanding Model Evaluator is a remote evaluation track...  ...modeling team can use to retrain. AI data reviewers help turn...  ...unsafe response with the correct policy category and severity. Audit...  ..., content moderation, or trust & safety review. Experience with inter... 
    Policy
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    4 days ago
  •  ...rigorous research to deepen our understanding of complex Trust & Safety challenges and inform better policy, product, and operational decisions. This role...  ...research, and domain expertise to uncover emerging risks, evaluate safety interventions, and translate insights into... 
    Policy

    TikTok

    New York, NY
    2 days ago
  • $80 - $120 per hour

    Product management / roadmap / PRD Evaluator is a remote evaluation track...  ...team can use to retrain. AI data reviewers help turn...  ...unsafe response with the correct policy category and severity. Audit...  ...linguistics, content moderation, or trust & safety review. Experience with inter... 
    Policy
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    4 days ago
  • As a Policy Manager, Search Policy in our Trust & Safety team, you will have the opportunity to work with teams of product...  ...search indexing, ranking, and AI features (e.g., AI Overviews, Snippets...  ...behaviors.- Design benchmarks and evaluation suites to measure policy... 
    Policy
    Immediate start

    TikTok

    New York, NY
    3 days ago
  •  ...is seeking a Senior Manager, Scaled Detection & Enforcement, Trust & Safety, to lead a distributed team focused on mitigating violative content...  ...strategy, optimize labeling operations, and collaborate with Policy, Product, Engineering, Analytics and Vendor Operations to... 
    Policy

    Etsy, Inc.

    New York, NY
    2 days ago
  • Kindred is seeking an experienced Trust & Safety leader to govern risk across our marketplace. You will own T&S policies, enforcement standards, and the appeals framework while analyzing trends and guiding product decisions. You will detect fraud through pattern analysis... 
    Policy

    Kindred

    New York, NY
    4 days ago
  • $65 - $70 per hour

     ...Biology Expert (PhD) — AI Safety is a remote red-team track for stress...  ...patch the gap. Adversarial evaluation is how AuraOne hardens AI models...  ...weakness classes (jailbreak, policy bypass, prompt injection) for...  ...security, AppSec, or trust & safety operations. Experience... 
    Policy
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    4 days ago
  • The Trust & Safety team at TikTok ensures our global community is safe and empowered to create...  ...applications. The Youth Safety & Well-being Policy team is at the center of this mission,...  ...to develop and improve human and AI-driven moderation systems.Minimum Qualification... 
    Policy

    TikTok

    New York, NY
    3 days ago
  • The Trust & Safety team at TikTok ensures our global community is safe and empowered to create...  ...applications. The Youth Safety & Well-being Policy team is at the center of this mission,...  ...to develop and improve human and AI-driven moderation systems.Minimum Qualifications... 
    Policy
    Immediate start

    TikTok

    New York, NY
    2 days ago
  • AuraOne is seeking a remote Trust and Safety Policy AI Evaluator to stress-test AI systems against adversarial prompts. You will craft attack scenarios, document failures, and map each jailbreak to the violated rubric clause to help patch gaps. You will think like attackers... 
    Policy
    Remote job
    Hourly pay
    For contractors

    AI Trainer Jobs

    New York, NY
    4 days ago
  •  ...is seeking a Director of Security Compliance and Trust to lead governance, risk, compliance, audit, privacy and AI governance across our cloud-native SaaS platform....  ...turn compliance into a business enabler, writing policy, designing controls and guiding #J-18808-Ljbffr... 
    Policy

    Motive

    New York, NY
    1 day ago
  •  ...controls for reliability, safety, and cost control—...  ...tracing ( AWS X-Ray ), evaluation frameworks, toxic content...  ...validation checkpoints. Zero-Trust Security: Design rigid...  ...to accelerate agentic AI development across the...  ..., throttling policies, and request/response transformations... 
    Policy
    Contract work

    Kelly Science, Engineering, Technology & Telecom

    New York, NY
    13 hours ago
  •  ...Data analysis / quantitative readouts Evaluator in a remote contractor role. You will review...  ...answer with rationale, and tagging issues like hallucinations or policy violations. Strong attention to detail and async availability are valued. #J-18808-Ljbffr AI Trainer Jobs
    Policy
    Remote job
    For contractors

    AI Trainer Jobs

    New York, NY
    4 days ago
  • AI Trainer Jobs seeks a Legal contracts / diligence / redlines Evaluator for a remote contractor role evaluating AI outputs in contract law workflows. Reviewers assess citations, statutory reasoning, and policy adherence, flag risk, and document the corrected analysis... 
    Policy
    Remote job
    Contract work
    For contractors

    AI Trainer Jobs

    New York, NY
    4 days ago
  •  ...across the marketplace and turn signals into policy. Combine analysis and governance to keep...  .... We\'re Looking For: 8+ years in Trust & Safety or risk with strong SQL and data skills....  ...fostering a culture of high standards and inclusivity. #J-18808-Ljbffr Kindred.ai
    Policy
    Remote job

    Kindred.ai

    New York, NY
    4 days ago
  •  ...experienced Operations Manager III to oversee quality programs within Trust & Safety and Content Moderation operations. This remote role in New...  ...management, and driving cross-functional initiatives with Policy, Engineering, and Operations. The ideal candidate will have extensive... 
    Policy
    Remote work

    Russell Tobin

    New York, NY
    1 day ago
  • $9 - $30 per hour

     ...participates in E-Verify. Description About the Job The Testing Evaluator position is a remote, on demand, part-time position. The...  ...workplace. All employees are required to adhere to the Company's drug-free workplace policy. #J-18808-Ljbffr ALTA Language Services, Inc.
    Policy
    Full time
    Part time
    For contractors
    Currently hiring
    Work at office
    Remote work
    Flexible hours

    ALTA Language Services, Inc.

    New York, NY
    1 day ago
  • $20 - $30 per hour

    A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong... 
    Remote job
    Hourly pay

    Crossing Hurdles

    New York, NY
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Trust and Safety Policy AI Evaluator. Be the first to apply!