Operations Research Model Prompt Evaluator [Remote]
$60 - $80 per hourAuraOne Human Data
- Remote job
Operations Research Model Prompt Evaluator is a remote review track for evaluating AI outputs across operations research model prompt research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Operations Research Model Prompt research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current operations research model prompt research review methods, conventions, and prior work for Operations Research Model Prompt Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in operations research model prompt research review or a closely related field for Operations Research Model Prompt Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a operations research model prompt research review derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Operations Research Model Prompt research review
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
$60–$80 / hr
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- ...topline and streamline their operations. We specialize in providing tailored... ...: Linkedin Job Title: Evaluator - Political Science Location:... ...question posed in the user prompt. The goal is to assess quality... ...Stay updated with the latest research, guidelines, and advancements...OperationsContract workRemote workWorldwide
- ...Role Overview Help evaluate frontier AI systems’ ability... ...scenarios, assess model responses, and define... ...challenging single-turn prompts in radiological safety... ...is preferred. Operational health physics experience... ...Technical writing, published research, or expert witness...OperationsRemote work
- ...Refusal Preference Reward Model Evaluator is a remote red-team track for... ...systems against adversarial prompts. Reviewers craft attack scenarios... ...AI systems, security research, or adversarial ML work for... ..., AppSec, or trust & safety operations. Experience publishing or...OperationsRemote jobHourly payFor contractors10 hours per week
- ...network and facility operations. Bitdeer also offers advanced... ...product built on a model is bounded by what it... ...This role builds the evaluation and decision systems... ...consumption. You will research and prototype adaptive... ...calling, context windows, prompt caching, reasoning...OperationsFull time
- ...Annotator—Product Management & Marketing to remotely review evaluation prompts and responses against the company's quality rubric. You will... ...outputs, label edge cases, and provide structured feedback the modeling team can use to retrain. As a contractor, you will assess...SuggestedRemote jobFor contractors
- AI Trainer Jobs is seeking an Illustration Quality Evaluator for a remote contractor role. Review illustration quality evaluation prompts and responses against a published rubric, compare paired outputs, and provide structured feedback to support retraining efforts. You...Remote jobPart timeFor contractors
- AI Trainer Jobs is seeking a remote contractor to evaluate people ops / recruiting prompts and responses against a evolving quality rubric. You will compare model outputs, label edge cases, and provide structured feedback for retraining. Responsibilities include evaluating...Remote jobFor contractors
- Surgical Planning Safety Evaluator is a remote evaluation track for reviewing surgical planning safety evaluation prompts and responses against AuraOne's quality rubric. Reviewers... ...write the kind of structured feedback the modeling team can use to retrain. AI data...Remote job
- Receipt and Invoice Understanding Model Evaluator is a remote evaluation track for reviewing receipt and invoice understanding model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind...Hourly payFor contractorsRemote work10 hours per week
- ...Certified Evaluator The Certified Evaluator delivers... ...mobile appointments, researches and evaluates merchandise... ...details, and promptly communicate changes or... ...Accurately record brand, model, serial number, condition... ...needs or damage. Daily Operations and Team...OperationsMinimum wageTemporary workWork at officeLocal areaRelocation
$65 - $70 per hour
...and Safety Policy AI Evaluator is a remote red-team track... ...against adversarial prompts. Reviewers craft... ...how AuraOne hardens AI models before they ship to customers... ...AI systems, security research, or adversarial ML... ...AppSec, or trust & safety operations. Experience publishing...OperationsHourly payFor contractorsRemote work10 hours per week$173k
...ManagementCompany: CitiCitibank, N.A. seeks a Model Validation 2nd LOD Lead Analyst for its... ...processes, aiming to increase operational efficiency for the organization. Developed... ...Transformer Architectures, RAG Pipelines, Prompt Engineering, Fine-Tuning (LoRA, RLHF, PEFT...OperationsFull timeRemote work- ...profitable growth of the company. The Fit Model has an integral role in the product... ...patterns, and retrieve or ship packages. Operate within Agenda. Meet all assigned dates and... ...Fitting Schedules and Business Meetings promptly. Collaborate Effectively with Colleagues....OperationsContract workPart timeCasual workShift work
$40 - $100 per hour
...About AI Training and Scientific Evaluation AI training is the human side... .... Specialists review model responses, test reasoning, identify... ...-by-step solutions Evaluate prompts, model answers, and technical... ...scenarios, calculations, and research-style reasoning tasks. Focus...Hourly payContract workPart timeFor contractorsRemote work$90 - $120 per hour
...Create benchmark-quality responses for future model outputs and document the context and... ...educational statistics materials, such as research articles, curricula, industry reports, or... ...each week. Availability to begin promptly is expected. Selected candidates should be...Hourly payFor contractorsRemote work$125k - $150k
...ECS is seeking an AI Model Engineer to work in a... ...with a track record of evaluating, experimenting with, and... ...technologies into operational use. This role requires... ...execution of cutting-edge research programs designed to... ...agentic workflow, and prompt engineeringExperience...OperationsContract workWork at officeRemote work- ...AI Labs seeks PhD and academic experts to support AI research projects remotely. You will evaluate model outputs across technical and humanities topics,... ...and depth. As a contractor, you will create reference prompts, critiques, and quality benchmarks, and contribute to...Remote jobFor contractors
$50 per hour
Combinatorics Model Evaluator is a remote review track for evaluating AI outputs across combinatorics model research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method...Hourly payFor contractorsRemote work$15 - $20 per hour
...creative and technical talent with leading AI research labs. Headquartered in San Francisco,... ...tools. Generate high-quality human evaluation data by identifying response strengths,... ..., and completeness of responses. Ensure model responses align with expected conversational...Contract workSummer workRemote work- ...Multimodal Hallucination Model Evaluator is a remote evaluation track for reviewing multimodal hallucination model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback...Remote jobHourly payFor contractors10 hours per week
- YO AI Labs is seeking Humanities Evaluation Specialists for a remote contract to support an AI training project. You will research, analyze, and craft challenging humanities questions... ...experience is required. You will design prompts that require deep interpretation and...Remote jobContract work
- ...a PhD & Academic Expert to support AI research projects remotely. You will apply your subject-matter expertise to evaluate and improve model responses across technical and humanities... ...disciplines. You will develop expert prompts, reference answers, and critiques, while...Remote job
$50 - $100 per hour
...engineering and problem-solving expertise to code generation and model evaluation work that helps improve how next-generation AI systems learn,... ..., with preference for experts who are available to begin promptly. Participation is subject to selection for the project....Hourly payContract workFor contractorsRemote work$37.5 per hour
...originals. In this role you will be responsible for evaluating advertising copy generated by a large language model (LLM) to ensure that it meets their benchmark in... ...expertise in strategy, design, execution and operations unlocks business value through a range of...OperationsContract workTemporary workRemote work- ...computational problem solving to improve and evaluate large language models. You will design rigorous math... ...customers Accelerate frontier AI research by contributing high quality data and... ...topics. Design exact, closed ended prompts and produce reliable Python solutions...Contract workFor contractorsFreelanceRemote work
- AuraOne is seeking a Latin Bilingual Expert for a remote evaluation track. Reviewers compare paired outputs to a quality rubric, label edge cases, and generate structured feedback to retrain models. Responsibilities include evaluating outputs against rubrics, tagging issues...For contractorsRemote work10 hours per week
- ...Research Interns at Applied Intuition Applied Intuition, Inc. is powering the future... ...three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the... ...on pretraining world-action foundation model with various world modalities including...Full timeFor contractorsFor subcontractorCasual workInternshipWork at officeImmediate startRemote workDay shift
- ...Role Overview Help evaluate how advanced AI systems handle sensitive... ...misuse potential, ensuring models remain useful for routine professional... ...challenging single-turn prompts from your domain and classify... ...writing ability. Published research, prior technical writing, or...For contractorsRemote work
$65 - $75 per unit
...analytical chemistry expertise to evaluate how AI systems handle... ...could enable harm, ensuring models provide useful answers when appropriate... ...single-turn chemistry prompts classified as benign, dual-use... ...technical writing ability. Published research, prior technical writing, or...Remote work$16 - $20 per hour
...Research Evaluator Position The research team at Penn State Ross and Carol Nese College of Nursing is hiring part-time research evaluators for projects focused on dementia care in assisted living settings. The research evaluator will assist with in-person recruitment...Hourly payPart timeSummer work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Operations Research Model Prompt Evaluator [Remote]. Be the first to apply!



