Refusal Preference Reward Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap.
Why this role matters
Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them.
Responsibilities
- Design adversarial prompts that probe known weakness classes (jailbreak, policy bypass, prompt injection) for Refusal Preference Reward Model Evaluator assignments.
- Document every successful attack with reproduction steps and the policy clause it violated.
- Score model defenses across single-turn and multi-turn conversations.
- Triage emerging attack vectors and route them to the safety team with severity ratings.
- Maintain a personal library of attack patterns and propose new red-team rubrics.
- Calibrate against the broader red-team cohort to keep coverage and severity consistent.
Qualifications
- Demonstrated experience red-teaming AI systems, security research, or adversarial ML work for Refusal Preference Reward Model Evaluator work.
- Strong written communication — your reports become the patch ticket.
- Comfort working in policy-grey areas with clear documentation of what was attempted and why.
- Familiarity with prompt-injection, jailbreak, and policy-bypass taxonomies.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Construct a 5-turn adversarial conversation that bypasses a specific policy clause and write up the patch ticket.
- Score a model's defenses against a known jailbreak pattern across 20 variants.
- Propose a new red-team rubric category after spotting an emerging attack vector.
- Reproduce a failure another reviewer reported and confirm the severity tag.
Nice to have
- Background in offensive security, AppSec, or trust & safety operations.
- Experience publishing or reproducing public adversarial-ML research.
- Multilingual fluency for cross-language attack testing.
Skills
- Adversarial prompting
- Red-team analysis
- Policy taxonomy
- Failure documentation
- Refusal Preference Reward Model evaluation
- Preference ranking
- RLHF
- Rater calibration
- Refusal
- Preference
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Conversational Preference Reward Model Evaluator is a remote evaluation track for reviewing conversational preference reward model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind...SuggestedRemote jobHourly payFor contractors10 hours per week
$20 per hour
...external tools. Generate high-quality human evaluation data by identifying response strengths,... ..., and completeness of responses. Ensure model responses align with expected... ...requiring structured analytical thinking Preferred Experience with RLHF, model evaluation,...SuggestedRemote jobContract workPart timeSummer work$178.5k - $257.83k
...Scientist, Transgenic Model and TechnologyLocation... ...efficacy studies• Identify, evaluate, and implement new... ..., neuroscience preferred)Leadership & Strategic... ...thoughtful, well-crafted rewards package that recognizes... ...information (including the refusal to submit to genetic...SuggestedFull timeContract workWork at officeRemote workFlexible hours- ...Program Analytics & Insights Evaluator IIon the project, you will be... ...evaluation frameworks, logic models, research questions, and evaluation... ...or at any time in the future Preferred: Experience using qualitative... ...Deloitte one of the most rewarding places to work. Our...SuggestedLocal area
$182.4k - $273.6k
...deliver scalable, production ready models and AI driven decision... ...shared expectations for quality, evaluation rigor, and production... ...higher degree. Master's or Ph.D. preferred in Machine Learning, Applied... ...package for employees. Other rewards may include short-term or annual...Temporary workWork at officeRemote workShift work3 days per week$60 - $90 per hour
...Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract... ...learning ideas, focusing on reward functions and training behavior. Evaluate... ...Proficiency in Python and Git . Preferred Basic understanding of...Full timeContract workSummer workRemote work$46.83k
...Education Certification Evaluator Are you an analytically minded professional eager... ...resolution efforts Apply today to begin a rewarding, yet challenging career as you... ...(51 Pa. C.S. 7103) provides employment preference for qualified veterans for appointment...Full timePart timeWork at officeLocal areaRemote workWork from homeMonday to Friday$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours$219k - $351k
...memory-bandwidth business. As models scale past what any single... ...LabLocation: Daily onsite presence preferred at our San Jose office/... ...incentive opportunities that reward employees based on individual... ...to ensure every candidate is evaluated fairly and holistically.Recruiting...Work at officeRemote workFlexible hours$172.43k - $230.95k
...Senior Software Engineer for the AI Model Lifecycle team will play a... ...reinforcement learning pipelines (e.g., preference optimization, policy optimization, reward modeling). Dataset, model, and... ...management: versioning, lineage, evaluation, and reproducible fine-tuning at...Full timeTemporary work- ...ServicesProfessional (Substance Use Disorder Evaluator) to provide substance use... ...and practices, patient preferred languages, health literacy,... ...or was unable to consent or refuse; or, Has been civilly or... ...: The team works in a hybrid model, with 3-4 days in the office...Contract workTemporary workWork experience placementWork at officeRemote workTrial periodMonday to FridayFlexible hoursNight shift
- A luxury brand evaluation company is seeking a Luxury Brand Evaluator to assess customer experiences with premium brands. The role offers flexibility in choosing assignments and entails evaluating services in-store or online. Ideal candidates are detail-oriented, observant...Flexible hours
$96.13k - $144.19k
...will support TD Bank's treasury model development team in... ...office; "NOT" Remote eligible) Preferred Qualifications: ~ Quantitative... ...related information to develop and evaluate options and implement... ...and communities. Our Total Rewards Package Our Total Rewards...Full timeWork experience placementWork at officeLocal areaFlexible hours- ...Harmful Instruction Refusal Risk Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft... ...matters Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like...Remote jobHourly payFor contractors10 hours per week
- ...vehicle design. As a Senior Virtual Hardware Model Engineer, you have the unique... ...related to software/hardware connectivity. Preferred experience in translating physical phenomena... ...your ambitions. Learn how GM supports a rewarding career that rewards you personally by...Full timeLocal areaWork from homeRelocation package
$265k - $360k
...is seeking an experienced and motivated Model Sales and Strategy Lead to join our U.S.... ...Bachelor’s degree required; MBA and or CFA preferred.Knowledge of fixed income, multi-sector... ...a total compensation approach when rewarding employees which includes a base salary and...Full timeHome officeFlexible hours$110.6k - $178k
Principal Model Based Design Engineer - HybridThis is a hybrid position that requires onsite... ...Control Systems, or a related field (PhD preferred).10+ years of experience in control... ...that salary is only one component of total rewards at Stanley Black & Decker. The salary range...Full timeLocal area- ...Office Specialist Vi: Transcript Evaluator St. Edward's University of Austin, Texas invites... ...in higher education or related field. Preferred qualifications include a Bachelor's... ...The University offers an excellent total rewards package! Medical & Rx Coverage (HSA Available...Temporary workInternshipWork at officeImmediate startFlexible hours
$75.33k - $125.5k
...employees. Collaborative, in-office operating model Retirement program (401k and Pension)... ...or interest-rate model techniques preferred; Experience using machine learning techniques... ...Perks At FHLBank Chicago, we believe in rewarding our high performing workforce. We offer...Work experience placementWork at officeRemote work$107.9k - $195.05k
...currently has an opening for a cleared Model Based Systems Engineer on the Digital Infrastructure... ...tools (e.g. Jira and Confluence). Preferred Qualifications: Active TS/SCI... ...the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We're...For contractorsLocal areaImmediate startRemote work- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers grade differential reasoning, dosing logic, and guideline adherence; flag patient-safety issues; and document the corrected clinical reasoning...Remote jobHourly payFor contractors10 hours per week
- ...Lean Theorem Model Evaluator is a remote review track for evaluating AI outputs across lean theorem model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method...Remote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
$107.5k - $204.5k
...Will Do * Lead the enterprise common data model across conceptual, logical, and physical... ...program artifacts Qualifications We Prefer: Experience in aerospace, defense, or... ...sessions. Healthy You Incentives, wellness rewards program. Doctor on Demand, virtual...Full timeContract workTemporary workWork experience placementWork at officeRemote workFlexible hours$76.29k - $114.44k
...will support TD Bank's treasury model development team in... ...office; "NOT" Remote eligible) Preferred Qualifications: ~ Quantitative... ...related information to develop and evaluate options and implement... ...and communities. Our Total Rewards Package Our Total Rewards...Full timeWork experience placementWork at officeLocal areaFlexible hours$86.84k - $139.36k
...industry best practices. The Non-Model/End-User-Computing Tool (EUC)... ...lead, plan, implement, and evaluate program/project activities to... ...information with discretion Preferred Qualifications Non-Model/EUC... ...- and so will you. Our Total Rewards Package Our Total Rewards package...Local areaWork from homeFlexible hours- ...The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent... ...services setting as a member of a multi-disciplinary team, preferred. EDUCATION ~ Masters Degree in Social Work or Mental Health...
$120k - $140k
SENIOR MANAGER OF SALES/NEW MODEL REQUIRED: Automotive experience... ...equipment. · Monitor and evaluate warranty concerns to ensure that... ...arise. HOW YOU WILL BE REWARDED: · Medical, Dental, Vision... ...or related Engineering field, preferred · 15 years of automotive...Work at officeRemote workVisa sponsorshipNight shift- ...conversion. Qualifications Education and Training: Degree in a health-related field from an accredited college or university preferred. Current licensure in physical therapy or occupational therapy is required. Prior marketing and/or rehabilitation/LTACH experience...Full timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Refusal Preference Reward Model Evaluator [Remote]. Be the first to apply!




