Refusal Preference Reward Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap.
Why this role matters
Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them.
Responsibilities
- Design adversarial prompts that probe known weakness classes (jailbreak, policy bypass, prompt injection) for Refusal Preference Reward Model Evaluator assignments.
- Document every successful attack with reproduction steps and the policy clause it violated.
- Score model defenses across single-turn and multi-turn conversations.
- Triage emerging attack vectors and route them to the safety team with severity ratings.
- Maintain a personal library of attack patterns and propose new red-team rubrics.
- Calibrate against the broader red-team cohort to keep coverage and severity consistent.
Qualifications
- Demonstrated experience red-teaming AI systems, security research, or adversarial ML work for Refusal Preference Reward Model Evaluator work.
- Strong written communication — your reports become the patch ticket.
- Comfort working in policy-grey areas with clear documentation of what was attempted and why.
- Familiarity with prompt-injection, jailbreak, and policy-bypass taxonomies.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Construct a 5-turn adversarial conversation that bypasses a specific policy clause and write up the patch ticket.
- Score a model's defenses against a known jailbreak pattern across 20 variants.
- Propose a new red-team rubric category after spotting an emerging attack vector.
- Reproduce a failure another reviewer reported and confirm the severity tag.
Nice to have
- Background in offensive security, AppSec, or trust & safety operations.
- Experience publishing or reproducing public adversarial-ML research.
- Multilingual fluency for cross-language attack testing.
Skills
- Adversarial prompting
- Red-team analysis
- Policy taxonomy
- Failure documentation
- Refusal Preference Reward Model evaluation
- Preference ranking
- RLHF
- Rater calibration
- Refusal
- Preference
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Style Preference Reward Model Evaluator is a remote evaluation track for reviewing style preference reward model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback...SuggestedRemote jobHourly payFor contractors10 hours per week
- AuraOne is seeking a remote Preference Dataset QA Reward Model Evaluator to review prompts and responses against our quality rubric. You will compare paired outputs, label edge cases, and write structured feedback for the modeling team to retrain. As a contractor, you...SuggestedRemote jobHourly payFor contractors
$50 per hour
Combinatorics Model Evaluator is a remote review track for evaluating AI outputs across combinatorics model research review reasoning, calculations... ...Creating a specialist profile records your experience and preferences. Starting role intake is a separate action that attaches...SuggestedHourly payFor contractorsRemote work$15 - $20 per hour
...external tools. Generate high-quality human evaluation data by identifying response strengths,... ..., and completeness of responses. Ensure model responses align with expected... ...requiring structured analytical thinking. Preferred Prior experience with RLHF, model evaluation...SuggestedContract workSummer workRemote work$40 per hour
Chart Understanding Model Evaluator is a remote evaluation track for reviewing chart understanding model evaluation prompts and responses... ...Creating a specialist profile records your experience and preferences. Starting role intake is a separate action that attaches this...Hourly payFor contractorsRemote work10 hours per week- Receipt and Invoice Understanding Model Evaluator is a remote evaluation track for reviewing receipt and invoice understanding model evaluation... ...Creating a specialist profile records your experience and preferences. Starting role intake is a separate action that attaches...Hourly payFor contractorsRemote work10 hours per week
$85 per hour
...Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud... ...infrastructure and reliability engineering solutions. Preferred ~ Experience supporting production-scale...Contract workSummer workRemote work- ...Workflow Annotator—Product Management & Marketing to remotely review evaluation prompts and responses against the company's quality rubric.... ..., label edge cases, and provide structured feedback the modeling team can use to retrain. As a contractor, you will assess frontier...Remote jobFor contractors
- Surgical Planning Safety Evaluator is a remote evaluation track for reviewing surgical planning safety evaluation prompts and responses... ...edge cases, and write the kind of structured feedback the modeling team can use to retrain. AI data reviewers help turn surgical...Remote job
- AI Trainer Jobs is seeking a remote contractor to evaluate people ops / recruiting prompts and responses against a evolving quality rubric. You will compare model outputs, label edge cases, and provide structured feedback for retraining. Responsibilities include evaluating...Remote jobFor contractors
- AI Trainer Jobs is seeking an Illustration Quality Evaluator for a remote contractor role. Review illustration quality evaluation prompts and responses against a published rubric, compare paired outputs, and provide structured feedback to support retraining efforts. You...Remote jobPart timeFor contractors
- AuraOne is seeking a Latin Bilingual Expert for a remote evaluation track. Reviewers compare paired outputs to a quality rubric, label edge cases, and generate structured feedback to retrain models. Responsibilities include evaluating outputs against rubrics, tagging issues...For contractorsRemote work10 hours per week
$150k - $175k
...for this role.Job Description:Preferred Qualifications:TD Securities... ...TDS) is seeking an experienced Model Governance professional to... ...expert (SME) responsible for evaluating the current state of model governance... ...effectively addresses and/or rewards performance in a timely...Full timeTemporary workLocal areaWork from homeFlexible hours$90k - $144.9k
Project Engineer, Model Based Design - Electric Motors - HybridThis is a Hybrid position... ...electrical engineering or a related field (PhD preferred).5+ years of experience in... ...that salary is only one component of total rewards at Stanley Black & Decker. The salary range...Full timeLocal area$219k - $351k
...memory-bandwidth business. As models scale past what any single... ...LabLocation: Daily onsite presence preferred at our San Jose office/... ...incentive opportunities that reward employees based on individual... ...to ensure every candidate is evaluated fairly and holistically.Recruiting...Work at officeRemote workFlexible hours$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours- A luxury brand evaluation company is seeking a Luxury Brand Evaluator to assess customer experiences with premium brands. The role offers flexibility in choosing assignments and entails evaluating services in-store or online. Ideal candidates are detail-oriented, observant...Flexible hours
- ...Harmful Instruction Refusal Risk Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft... ...matters Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like...Remote jobHourly payFor contractors10 hours per week
$265k - $360k
...is seeking an experienced and motivated Model Sales and Strategy Lead to join our U.S.... ...Bachelor’s degree required; MBA and or CFA preferred.Knowledge of fixed income, multi-sector... ...a total compensation approach when rewarding employees which includes a base salary and...Full timeHome officeFlexible hours$35.81 - $42 per hour
...Health is looking for a Part-time PASRR Evaluator to join our growing team. Job Summary: Join... ...mental illness, or related conditions Preferred Qualifications Knowledge of Federal... ...Benefits Benefits are a key component of your rewards package. Our benefits are designed to...Contract workPart timeWork at officeLocal areaRemote workWork from home- ...DescriptionThe RoleAs a Vehicle Energy Model and Toolchain Development Engineer, you... ...customerWhat Will Give You A Competitive Edge (Preferred Qualifications)7+ years of relevant... ...your ambitions. Learn how GM supports a rewarding career that rewards you personally by visiting...Full timeLocal areaWork from home
- ...complete answer and when it requires a refusal, including scenarios involving isotopes,... ...as benign, dual-use, or adversarial. Evaluate AI responses against a defined policy standard... ...users. Red-teaming experience is preferred. Technical writing, published research,...For contractorsRemote work
$180k - $240k
...Senior/Staff Software Engineer (Platform and Execution Model) Seattle, WA (Preferred) or McLean, VA or Remote (USA) About Us Co-founded... ...observability: distributed tracing, structured logs, metrics, and evaluation hooks; build an "explainable trail" of agent actions...Remote jobFull timeShift work- ...Theorem Proving Model Evaluator is a remote review track for evaluating AI outputs across theorem proving model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct...Remote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers grade differential reasoning, dosing logic, and guideline adherence; flag patient-safety issues; and document the corrected clinical reasoning...Remote jobHourly payFor contractors10 hours per week
- .../Education/Certification): High School Diploma or higher. Preferred but not required a minimum of one year of successful experience... ...skills, and abilities required. The district shall not fail or refuse to hire or discharge any individual, or otherwise discriminate...Work at office
$120k - $140k
SENIOR MANAGER OF SALES/NEW MODEL REQUIRED: Automotive experience... ...equipment. · Monitor and evaluate warranty concerns to ensure that... ...arise. HOW YOU WILL BE REWARDED: · Medical, Dental, Vision... ...or related Engineering field, preferred · 15 years of automotive...Work at officeRemote workVisa sponsorshipNight shift$60 - $90 per hour
...Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract... ...reinforcement learning ideas, focusing on reward functions and training behavior.... ...Proficiency in Python and Git . Preferred Basic understanding of reinforcement...Contract workSummer workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Refusal Preference Reward Model Evaluator [Remote]. Be the first to apply!



