Refusal Preference Reward Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap.
Why this role matters
Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them.
Responsibilities
- Design adversarial prompts that probe known weakness classes (jailbreak, policy bypass, prompt injection) for Refusal Preference Reward Model Evaluator assignments.
- Document every successful attack with reproduction steps and the policy clause it violated.
- Score model defenses across single-turn and multi-turn conversations.
- Triage emerging attack vectors and route them to the safety team with severity ratings.
- Maintain a personal library of attack patterns and propose new red-team rubrics.
- Calibrate against the broader red-team cohort to keep coverage and severity consistent.
Qualifications
- Demonstrated experience red-teaming AI systems, security research, or adversarial ML work for Refusal Preference Reward Model Evaluator work.
- Strong written communication — your reports become the patch ticket.
- Comfort working in policy-grey areas with clear documentation of what was attempted and why.
- Familiarity with prompt-injection, jailbreak, and policy-bypass taxonomies.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Construct a 5-turn adversarial conversation that bypasses a specific policy clause and write up the patch ticket.
- Score a model's defenses against a known jailbreak pattern across 20 variants.
- Propose a new red-team rubric category after spotting an emerging attack vector.
- Reproduce a failure another reviewer reported and confirm the severity tag.
Nice to have
- Background in offensive security, AppSec, or trust & safety operations.
- Experience publishing or reproducing public adversarial-ML research.
- Multilingual fluency for cross-language attack testing.
Skills
- Adversarial prompting
- Red-team analysis
- Policy taxonomy
- Failure documentation
- Refusal Preference Reward Model evaluation
- Preference ranking
- RLHF
- Rater calibration
- Refusal
- Preference
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- ...Preference Taxonomy Reward Model Evaluator is a remote evaluation track for reviewing preference taxonomy reward model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...SuggestedRemote jobHourly payFor contractors10 hours per week
$20 per hour
...external tools. Generate high-quality human evaluation data by identifying response strengths,... ...and completeness of responses. Ensure model responses align with expected... ...requiring structured analytical thinking Preferred Experience with RLHF, model evaluation,...SuggestedRemote jobPart timeSummer work$178.5k - $257.83k
...Scientist, Transgenic Model and TechnologyLocation... ...efficacy studies• Identify, evaluate, and implement new... ..., neuroscience preferred)Leadership & Strategic... ...thoughtful, well-crafted rewards package that recognizes... ...information (including the refusal to submit to genetic...SuggestedFull timeContract workWork at officeRemote workFlexible hours- Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts...SuggestedRemote job
$182.4k - $273.6k
...deliver scalable, production ready models and AI driven decision... ...shared expectations for quality, evaluation rigor, and production... ...higher degree. Master's or Ph.D. preferred in Machine Learning, Applied... ...package for employees. Other rewards may include short-term or annual...Temporary workWork at officeRemote workShift work3 days per week$19.25 per hour
...operate a company vehicle. Assist in Fruit Evaluation Management team duties Monitor crop... ...with a specific emphasis in agriculture preferably within a production/manufacturing setting... ...engaging employee community groups cash rewards for healthy habits and fitness reimbursements...Seasonal workWork at officeLocal areaWorldwideFlexible hoursAfternoon shift$60 - $90 per hour
...Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract... ...learning ideas, focusing on reward functions and training behavior. Evaluate... ...Proficiency in Python and Git . Preferred Basic understanding of...Full timeContract workSummer workRemote work- ...Audit function! It is our vision to be a preferred advisor to the business by building... ...more about you!Support Internal Audit’s evaluation of model and artificial intelligence (AI) risk... ...appreciation for our teams, who are rewarded with highly competitive pay and generous...InternshipMonday to Friday
$50 - $60 per hour
A technology company specializing in AI and finance is seeking a Wealth Advisor to help train AI models. In this independent contract role, you will measure the effectiveness of AI chatbots by solving complex financial problems. The ideal candidate should be fluent in...Remote jobHourly payContract work$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours- A luxury brand evaluation company is seeking a Luxury Brand Evaluator to assess customer experiences with premium brands. The role offers flexibility in choosing assignments and entails evaluating services in-store or online. Ideal candidates are detail-oriented, observant...Flexible hours
$36 - $40 per hour
...Acentra Health is looking for a Clinical Evaluator - PRN to join our growing team Job... ...commonly requiring long-term care placement. Preferred Qualifications 2 years of experience... ...Benefits Benefits are a key component of your rewards package. Our benefits are designed to...Hourly payContract workReliefWork at officeLocal areaRemote workWork from homeHome office$265k - $360k
...is seeking an experienced and motivated Model Sales and Strategy Lead to join our U.S.... ...Bachelor’s degree required; MBA and or CFA preferred.Knowledge of fixed income, multi-sector... ...a total compensation approach when rewarding employees which includes a base salary and...Full timeHome officeFlexible hours$110.6k - $178k
Principal Model Based Design Engineer - HybridThis is a hybrid position that requires onsite... ...Control Systems, or a related field (PhD preferred).10+ years of experience in control... ...that salary is only one component of total rewards at Stanley Black & Decker. The salary range...Full timeLocal area- ...vehicle design. As a Senior Virtual Hardware Model Engineer, you have the unique... ...related to software/hardware connectivity. Preferred experience in translating physical phenomena... ...your ambitions. Learn how GM supports a rewarding career that rewards you personally by...Full timeLocal areaWork from homeRelocation package
- ...Harmful Instruction Refusal Risk Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft... ...matters Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like...Remote jobHourly payFor contractors10 hours per week
$75.33k - $125.5k
...Model Risk Analyst At the Federal Home Loan Bank of Chicago, employees come first... ...processes or interest-rate model techniques preferred; Experience using machine learning... ...preferred At FHLBank Chicago, we believe in rewarding our high performing workforce. We offer...Work experience placementWork at officeRemote work$100k - $150k
...Large Language Model Specialist - Remote Bright Vision... ...language models across supervised, preference-based, and reinforcement... ...dataset construction, rigorous evaluation methodology, and the... ...probes. Implement safety, refusal, and policy evaluations to track...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$107.9k - $195.05k
...currently has an opening for a cleared Model Based Systems Engineer on the Digital Infrastructure... ...tools (e.g. Jira and Confluence). Preferred Qualifications: Active TS/SCI... ...the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We're...For contractorsLocal areaImmediate startRemote work$131.3k - $237.35k
...systems and our strategic systems. This Model-Based Systems Engineer will join a team... ...clearance or ability to obtain. Preferred Qualifications Preference shown to candidates... ...the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We'...Local areaImmediate startRemote workWorldwide$20 per hour
Feedinkoo is looking for a Web Developer/Designer to enhance AI models by evaluating design work, including interfaces and user experiences. This role involves reviewing AI‑generated visuals and providing feedback to improve users' experience with AI tools. Working remotely...Remote job- ...Number Theory Model Evaluator is a remote review track for evaluating AI outputs across number theory model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method...Remote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers grade differential reasoning, dosing logic, and guideline adherence; flag patient-safety issues; and document the corrected clinical reasoning...Remote jobHourly payFor contractors10 hours per week
$120k - $140k
SENIOR MANAGER OF SALES/NEW MODEL REQUIRED: Automotive experience... ...equipment. · Monitor and evaluate warranty concerns to ensure that... ...arise. HOW YOU WILL BE REWARDED: · Medical, Dental, Vision... ...or related Engineering field, preferred · 15 years of automotive...Work at officeRemote workVisa sponsorshipNight shift$60 - $90 per hour
...multi-step machine learning evaluation tasks for a leading generative... ...to determine where frontier models succeed or fail. Typical tasks... ...example modifying how an RL reward is computed and specifying success... ...and policy training is preferred. Prior experience in AI training...Hourly payFull timeFreelanceRemote work- ...The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent... ...services setting as a member of a multi-disciplinary team, preferred. EDUCATION ~ Masters Degree in Social Work or Mental Health...
$135.6k - $237.4k
Role Description The Director, Model Engineering & Operations is... ...monitoring practices. ~Guide the evaluation and responsible adoption of... ..., Data Science, Engineering preferred. ~Equivalent years of... ...Substantial and comprehensive total rewards package. Working...Full timeWork experience placementWork at office$117.3k - $226.9k
...control, and range systems. As the CAMEO Model-Based Systems Engineer - Space Force... ...Advanced degree in a STEM field is strongly preferred Directed multi‑disciplinary teams... ...competitive compensation package where you’ll be rewarded based on your performance and recognized...Full timeWork at officeImmediate startRemote workRelocation packageFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Refusal Preference Reward Model Evaluator [Remote]. Be the first to apply!





