Operations Research Model Prompt Evaluator [Remote]
$60 - $80 per hourAuraOne Human Data
- Remote job
Operations Research Model Prompt Evaluator is a remote review track for evaluating AI outputs across operations research model prompt research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document the correct method so the modeling team can train on it.
Why this role matters
Operations Research Model Prompt research review models live or die on whether their derivations actually hold up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a peer reviewer would — checking assumptions, reproducing key steps, and capturing the right method alongside the wrong one.
Responsibilities
- Review AI outputs against current operations research model prompt research review methods, conventions, and prior work for Operations Research Model Prompt Evaluator assignments.
- Reproduce or sanity-check key derivations, calculations, or experimental claims.
- Flag dimensional, methodological, and citation errors with structured severity tags.
- Capture the corrected reasoning or worked example so the modeling team can train on it.
- Adjudicate disputed answers against textbooks, papers, or community standards.
- Maintain reviewer-quality scores in inter-rater calibration cycles.
Qualifications
- Graduate-level training or equivalent applied experience in operations research model prompt research review or a closely related field for Operations Research Model Prompt Evaluator work.
- Hands-on experience publishing, teaching, or advising on the topic at a professional level.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that cites methods, papers, or worked examples.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Reproduce a operations research model prompt research review derivation from a model output and flag any algebraic or dimensional errors.
- Grade a model's literature summary against the cited papers and rate the citation quality.
- Adjudicate a disputed answer between two reviewers using textbook methods.
- Audit a 25-row batch for rubric consistency and report drift to the program lead.
Nice to have
- PhD, postdoc, or industry research experience in the topic area.
- Prior work reviewing AI-assisted research tooling and its failure modes.
- Multilingual fluency for non-English papers and corpora.
Skills
- Scientific reasoning
- Method validation
- Citation review
- Quantitative analysis
- Operations Research Model Prompt research review
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
$60–$80 / hr
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- ...topline and streamline their operations. We specialize in providing tailored... ...: Linkedin Job Title: Evaluator - Political Science Location:... ...question posed in the user prompt. The goal is to assess quality... ...Stay updated with the latest research, guidelines, and advancements...OperationsContract workRemote workWorldwide
- ...Refusal Preference Reward Model Evaluator is a remote red-team track for... ...systems against adversarial prompts. Reviewers craft attack scenarios... ...AI systems, security research, or adversarial ML work for... ..., AppSec, or trust & safety operations. Experience publishing or...OperationsRemote jobHourly payFor contractors10 hours per week
- Lynker Technologies Sea Ice Model Evaluations and Applications Scientist US-MD-Suitland Job ID: 2026-1639 Type: Full-Time # of Openings... ...) and the U.S. National Ice Center (USNIC), to support operational and research activities involving numerical sea ice forecast guidance...OperationsFull timeSeasonal workLocal areaRemote work
- ...experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for... ..., and data integrity, guiding model refinements and prompts for high-quality outputs. The role is remote, project...SuggestedRemote job
- ...this role and related opportunities across consulting, finance, product, software, AI, operations, market research, and growth-focused workstreams. THE ROLE As a Business Model Research Analyst, you will help identify promising markets, companies, products, operators...OperationsRemote workFlexible hours
- ...role and related opportunities. We work across consulting, finance, product, software, AI, operations, market research, and growth-focused workstreams. THE ROLE As a Business Model Research Analyst, you will help identify promising markets, companies, products,...OperationsRemote workFlexible hours
$300k - $320k
...Program Manager to lead our AI model evaluation initiatives across multiple... .... Working closely with our Research, Trust & Safety, Frontier... ...seen excellence at scale and operated in rapidly scaling, high-ambiguity... ...adoption. Have experience prompt engineering on language...OperationsWork at officeHome officeVisa sponsorshipRelocation package$20 per hour
...creative and technical talent with leading AI research labs. Headquartered in San Francisco,... ...tools. Generate high-quality human evaluation data by identifying response strengths,... ...and completeness of responses. Ensure model responses align with expected conversational...Remote jobPart timeSummer work- ...Visual Question Answering Model Evaluator is a remote evaluation track for reviewing visual question answering model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...Remote jobHourly payFor contractors10 hours per week
- ...Overview Drive the creation and evaluation of challenging STEM problems... ...and benchmark large language models. You will design multi-step... ..., and collaborate with researchers to build evaluation benchmarks... ...thinking when designing problem prompts and solutions. Feedback and...Contract workFor contractorsFreelanceRemote work
$60 - $80 per hour
...experienced biology professionals with AI research teams and companies. Members... ...domain expertise to help train and evaluate biology-focused AI models, design realistic tasks and deliverables... ...-related areas. Create tasks, prompts, and deliverables based on real world...Hourly payContract workImmediate startRemote work$15 - $20 per hour
...creative and technical talent with leading AI research labs. Headquartered in San Francisco,... ...tools . Generate high-quality human evaluation data by identifying response strengths,... ...completeness of responses. Ensure model responses align with expected conversational...Contract workSummer workRemote work$60 - $80 per hour
...world scientific input that helps models learn, reason, and perform. No... ...complex biochemical datasets, research findings, and protocols,... ...current biochemistry practices. Evaluate the scientific validity and clarity of biochemistry prompts, AI-generated outputs, and...Hourly payContract workRemote work$24 - $29 per hour
...Manager The Role: The Marketing Coordinator serves as a central operational partner across Brand Marketing, Creative Services, Editorial,... ...+ Insight Support Conduct ongoing competitive and market research across design, interiors, hospitality, luxury retail, and media...OperationsContract workRemote work- ...financial world.The roleSoFi is seeking a Fraud Model Developer to join our Fraud Model... ...team. In this role, you will develop, evaluate, and monitor machine learning models that... ...losses, minimize false positives, lower operational costs, and protect SoFi members. You will...OperationsRemote work
$60k - $100k
The ETF and Model Portfolio Business Development Associate will... ...independent advisors, home-office research teams, and key internal... ...Marketing, Legal, Compliance, Operations, and other internal stakeholders... ..., including the ability to evaluate market trends, sales data, product...OperationsFull timeTemporary workLocal areaHome officeFlexible hours$182.4k - $273.6k
...leaders to deliver scalable, production ready models and AI driven decision systems that... ...application of the Applied AI operating model, decision rights, delivery discipline... ...Reinforce shared expectations for quality, evaluation rigor, and production readiness.Provide...OperationsTemporary workWork at officeRemote workShift work3 days per week- ...adoption of AI by ensuring GenAI, model, and agentic systems are deployed, validated, and operated against published security... ...and run model validation and evaluation harnesses to test AI models and... ...guidance to engineering teams.Research, test, and pilot AI security tools...OperationsFull timePart timeWork at officeLocal areaWork from homeRelocationMonday to Thursday
- ...values in everything we do. NATURE AND SCOPE: The Lead Evaluator, Early Learning will provide high quality personalized and customized... ...to maintain a high level of knowledge about the current operations of Cognia by participating in on-going professional development...OperationsPart timeWork experience placementWork at officeLocal areaRemote work
$185k - $400k
...infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-training large-scale multimodal foundation models to advance our mission of...Remote work- ...company. We build cutting-edge foundation AI models and end-to-end products that are... ...for our customers. Cohere is a team of researchers, engineers, designers, and more, who are... ...Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling...Full timeWork at officeLocal areaRemote workHome office
$178.5k - $257.83k
...Distinguished Scientist, Transgenic Model and TechnologyLocation:... ...of in vivo and preclinical research programs while championing the... ...efficacy studies• Identify, evaluate, and implement new technologies... ...resource allocation and operational efficiencyOperational Excellence...Full timeContract workWork at officeRemote workFlexible hours- ...Research Intern Applied Intuition, Inc. is powering the future of physical AI. Founded... ...three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the... ...on pretraining world-action foundation model with various world modalities including...For contractorsFor subcontractorCasual workInternshipWork at officeImmediate startRemote workDay shift
- ...currently seeking a AI Foundational Model Engineer to join our team in... ...to deployment, monitoring, evaluation, rollback, and continuous... ...AI service components, APIs, prompts, retrieval logic, and observability... ...compliance, financial crime, operations, or enterprise technology...OperationsTemporary workWork at officeRemote workFlexible hours
- ...and customized by a team of experienced sellers, engineers, and researchers. Many of us worked on large-scale products including Llama,... ...collaboration with founders and execs Pioneer the training of new models that leverage both historical data and synthetic training data...Full time
$35 - $40 per hour
A leading evaluation organization seeks a Generalist Evaluator Expert for a remote, flexible contract role. This position involves designing prompts for language model evaluations, defining standards, and conducting assessments. Ideal candidates should have strong writing...Remote jobContract workFlexible hours- ...support the evolution of how Human Resources operates across the enterprise. This role will... ...strategies into a clear operating model, and ensure that work is tracked, measured... ...the Human Resources operating model by evaluating and enhancing governance processes, meeting...OperationsFull timeNight shiftWeekend work
- ...more about you!Support Internal Audit’s evaluation of model and artificial intelligence (AI) risk... ..., uncertainty, transparency, and operational risks associated with advanced analytical... ...storytelling and technical presentation skills Research Skills Interpersonal Skills Working...OperationsInternshipMonday to Friday
$170k - $250k
...enhancing the next layer: a production-grade capability to measure model performance and feed those insights back into how we build,... ...PartnershipPartner closely with Data & Analytics teams to establish a clear operating model, proactively collaborating with the business while...OperationsWork at officeLocal areaRemote workFlexible hours$16 - $20 per hour
...Research Evaluator Position The research team at Penn State Ross and Carol Nese College of Nursing is hiring part-time research evaluators for projects focused on dementia care in assisted living settings. The research evaluator will assist with in-person recruitment...Hourly payPart timeSummer work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Operations Research Model Prompt Evaluator [Remote]. Be the first to apply!




