Preference Dataset QA Reward Model Evaluator
AuraOne
Preference Dataset QA Reward Model Evaluator is a remote evaluation track for reviewing preference dataset qa reward model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain. Why this role matters AI data reviewers help turn preference dataset qa reward model evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data. Responsibilities Evaluate preference dataset qa reward model evaluation model outputs against a versioned rubric and assign severity tags for Preference Dataset QA Reward Model Evaluator assignments. Compare paired responses and pick the stronger answer with a written rationale. Label hallucinations, instruction-following failures, and unsafe content with structured tags. Capture ambiguous prompts and route them back to the program team for rubric updates. Maintain reviewer-quality scores by calibrating against gold-standard examples each week. Document recurring failure modes so the modeling team can target them in the next training run. Qualifications Prior evaluation, annotation, or human-rater experience on preference dataset qa reward model evaluation or adjacent content for Preference Dataset QA Reward Model Evaluator work. Comfort applying multi-page rubrics consistently across long batches. Clear written reasoning that names the issue and the rubric clause being applied. Strong attention to detail and the ability to flag when a prompt itself is the problem. Reliable async availability for at least 10 hours per week. Example tasks Compare two preference dataset qa reward model evaluation model responses to the same prompt and pick the stronger one with rationale. Tag an unsafe response with the correct policy category and severity. Audit a 50-row batch for rubric consistency and report drift to the program lead. Propose a rubric clarification after spotting a recurring failure mode. Nice to have Background in linguistics, content moderation, or trust & safety review. Experience with inter-rater agreement metrics and calibration cycles. Domain expertise that lets you spot subject-matter errors automated checks miss. Skills Model output evaluation Rubric-based annotation Severity tagging Inter-rater calibration Preference Dataset QA Reward Model evaluation Preference ranking RLHF Rater calibration Preference Dataset Work model Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US. Compensation Hourly rate confirmed after the interview process. Prior evaluation, annotation, or human-rater experience on preference dataset qa reward model evaluation or adjacent content for Preference Dataset QA Reward Model Evaluator work. Comfort applying multi-page rubrics consistently across long batches. Clear written reasoning that names the issue and the rubric clause being applied. Strong attention to detail and the ability to flag when a prompt itself is the problem. Reliable async availability for at least 10 hours per week. #J-18808-Ljbffr AuraOne
- AuraOne is seeking a remote Preference Dataset QA Reward Model Evaluator to assess model outputs against a versioned rubric and provide structured feedback. You will compare paired responses, label issues, and document edge cases for retraining the model. As an independent...SuggestedRemote jobFor contractors
- AuraOne is seeking a remote Pairwise Preference Reward Model Evaluator to review prompts and responses against our quality rubric. You will compare paired outputs, label edge cases, and provide structured feedback to help retrain the modeling team. This independent contractor...SuggestedRemote jobHourly payFor contractors10 hours per week
- AuraOne is seeking a remote Spoken Instruction Conversation Evaluator to review prompts and responses against its quality rubric. You will... ...edge cases, and provide structured feedback for retraining the model. As an independent contractor, you will evaluate model outputs,...SuggestedRemote jobHourly payContract workFor contractors
$40 per hour
...their team. This remote role involves training AI models by posing complex mathematical problems, evaluating outputs, and assessing the model's performance. Candidates... ...of mathematics. A relevant Master's or PhD is preferred but not required. The position offers the...SuggestedHourly payRemote work$20 per hour
...external tools. Generate high-quality human evaluation data by identifying response strengths,... ..., and completeness of responses. Ensure model responses align with expected... ...requiring structured analytical thinking Preferred Experience with RLHF, model evaluation,...SuggestedRemote jobContract workPart timeSummer work$50 - $100 per hour
DataAnnotation is seeking an experienced Legal Expert to help train AI models. You will tackle diverse legal problems, measure chatbot... .... Applicants must hold a J.D. and demonstrate expert English proficiency; US-based candidates preferred. #J-18808-Ljbffr DataAnnotationRemote jobHourly payFor contractorsFlexible hours- ...Senior Securitized Products Evaluator supports Evaluated... ...market data, prepayment models, yield curve movements,... ...-facing communication Preferred Skills Securitized product... ...waterfall structures QA process adherence... ...platform with best-in-class datasets, analytics, and technology...
- A leading AI development company is seeking experienced quantitative professionals for remote work evaluating AI-generated quantitative analysis. Ideal candidates will have a robust background in fields like data science, economics, or biostatistics, with at least 2 years...Remote work
$148.5k - $174.7k
...individual contributor to support our Model Development & Decision... ...and analysis: compile datasets, perform quality checks, and... ...validation processes is a plus.Preferred Skills· Comfort using Microsoft... ...approach to benefits and total rewards considers our team members’ whole...Full timeLocal area3 days per week- Productive Playhouse seeks AI Evaluators to support evaluating AI chatbots by interacting with models, assessing capabilities, safety, and usefulness. This is a project-based, task-based engagement with flexible hours and batch deliveries. Open to freelancers outside the...Remote jobFreelanceFlexible hours
- ...future team member for the role of SVP - Model Risk Management AI, Wealth and Investment... ...following: Advanced degree (Master’s or PhD preferred) in a quantitative field such as... ...Just Capital and CNBC, 2025Our Benefits and Rewards:BNY offers highly competitive compensation...WorldwideFlexible hours
$151k - $251.6k
LSEG's Evaluated Pricing Service provides end-of-day valuations across... ...trading activity, cash flow modeling, and relevant market news.... ...with daily production deadlines.Preferred QualificationsExperience with... ...working with large financial datasets and valuation model...Full timePart timeInternship- ...Product Manager in C360 - World Model, you will be the hands-on... ...vocabulary that keeps every team, dataset, and AI agent describing the... ...senior business stakeholders.Preferred qualifications, capabilities,... ...We offer a competitive total rewards package including base salary...Contract work
- ...member for the role of SVP - Model Risk Management to join our Model... ...’s degree required, PhD preferred.5-10 years of experience in model... ...skills, with the ability to evaluate complex model frameworks,... ...and CNBC, 2025Our Benefits and Rewards:BNY offers highly competitive...WorldwideFlexible hours
- A leading AI training company is seeking medical experts to evaluate AI chatbots' responses to complex healthcare problems in a REMOTE position. Applicants need to hold a medical degree or be in-progress towards one. Responsibilities include ensuring medical accuracy of...Remote jobHourly pay
$119k - $218.3k
...Enterprise Strategy, Risk and Operating Model Design Enterprise Operations & Risk Ready... ..., now or at any time in the future. Preferred qualifications Candidates possessing one... ...challenges. This makes Deloitte one of the most rewarding places to work. Our purpose...Contract workWork at office$212.6k - $250.1k
...into an integrated operating model for the Office of the CFO.As... ...location strategy, provider evaluation, transition planning, cutover... ...client and engagement needs.Preferred qualificationsExperience in a... ...Information on our competitive total rewards package, including our bonus...Work at officeLocal areaImmediate startFlexible hours$85 per hour
...Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud... ...infrastructure and reliability engineering solutions. Preferred ~ Experience supporting production-scale...Contract workSummer workRemote work$25 - $35 per hour
...Expert (Real Estate and Leasing) to work part-time and remotely. Responsibilities include writing prompts for real estate scenarios, evaluating AI outputs, and providing evidence-based evaluations. The ideal candidate has expertise in real estate or leasing, attention to...Remote jobHourly payPart time- AuraOne is seeking a Unit Test Generation Evaluation Specialist to remotely support evaluation workflows. You will review AI outputs, grade... ...structured rubrics, flag risks, and document next steps for model training. This contractor role requires reliable async availability...Remote jobHourly payFor contractors10 hours per week
$120 per hour
...Position: Document/deck production QA Evaluator Type: Contract Compensation:... ...structured written feedback to improve AI model outputs. Collaborate with AI research... ...Workspace , especially Slides . Preferred ~ Advanced degree ( Master's or higher...Remote jobContract workSummer workWork at office$207k - $301k
...both human-powered and Large Language Model (LLM)-powered automated evaluation systems to assess model performance... ...across different models and datasets.Provide actionable insights from evaluations... ...(LLM) interfaces into workflows.Preferred qualifications:Master’s degree or...$20 per hour
Feedinkoo is looking for a Web Developer/Designer to enhance AI models by evaluating design work, including interfaces and user experiences. This role involves reviewing AI‑generated visuals and providing feedback to improve users' experience with AI tools. Working remotely...Remote job- ...responsible for building and enhancing models that inform wind-down and business packaging... ...and coordinating inputs from others Preferred Qualifications, Capabilities, And Skills... ...management. We offer a competitive total rewards package including base salary determined...Work at officeVisa sponsorship
$101k - $203k
...Consulting practice and lead model validation and/or internal audit... ...R, Alteryx, or similar tools.Evaluate whether model validation and... ...as needed (estimated <30%).Preferred:Professional certification relevant... .... Learn more about our total rewards at .All applicants will...Full timeWork experience placementInternshipLocal area- AuraOne is seeking a Remote Social Engineering Security Evaluator to review prompts and responses against our quality rubric, producing... ...to support the security evaluation program. You will evaluate model outputs, tag issues, and collaborate with the team to improve training...Remote jobFor contractors
- ...clients globally by providing them evaluated pricing on over two million... ...of stochastic calculus, main models used within derivatives... ...adherence and quality control Preferred qualifications, capabilities,... ...We offer a competitive total rewards package including base salary...
- AuraOne is seeking a remote Probability Model Evaluator to review prompts and model outputs against a quality rubric. You will compare paired responses, tag edge cases, and provide structured feedback to help retrain the system. The role emphasizes clear reasoning, attention...Remote jobFor contractorsFlexible hours
- In-House Fit Model (Women's Size 8 / Medium)About the RoleWe are looking for a reliable... ...merchandising, and sales teams to help evaluate garment fit, comfort, and overall wearability... ..., samples, and future product launches.Preferred MeasurementsTo align with our...Full time
$65k - $179.4k
...maintaining Consumer and Commercial Models that support our retail and... ...experience in handling large datasets and deriving meaningful... ...modeling, credit cards experience preferred Experience with machine... ...models and assesses model risks Evaluates identified model risks and...Full timeTemporary workPart timeWork experience placement
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Preference Dataset QA Reward Model Evaluator. Be the first to apply!
- program evaluator New York, NY
- work from home web search evaluator New York, NY
- clinical evaluator New York, NY
- evaluator New York, NY
- ai evaluator New York, NY
- work from home social media evaluator New York, NY
- quality evaluator New York, NY
- education evaluator New York, NY
- social media evaluator New York, NY
- qa intern New York, NY

