Secure Code Review AI Evaluator [Remote]
AuraOne Human Data
- Remote job
Secure Code Review AI Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap.
Why this role matters
Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them.
Responsibilities
- Design adversarial prompts that probe known weakness classes (jailbreak, policy bypass, prompt injection) for Secure Code Review AI Evaluator assignments.
- Document every successful attack with reproduction steps and the policy clause it violated.
- Score model defenses across single-turn and multi-turn conversations.
- Triage emerging attack vectors and route them to the safety team with severity ratings.
- Maintain a personal library of attack patterns and propose new red-team rubrics.
- Calibrate against the broader red-team cohort to keep coverage and severity consistent.
Qualifications
- Demonstrated experience red-teaming AI systems, security research, or adversarial ML work for Secure Code Review AI Evaluator work.
- Strong written communication — your reports become the patch ticket.
- Comfort working in policy-grey areas with clear documentation of what was attempted and why.
- Familiarity with prompt-injection, jailbreak, and policy-bypass taxonomies.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Construct a 5-turn adversarial conversation that bypasses a specific policy clause and write up the patch ticket.
- Score a model's defenses against a known jailbreak pattern across 20 variants.
- Propose a new red-team rubric category after spotting an emerging attack vector.
- Reproduce a failure another reviewer reported and confirm the severity tag.
Nice to have
- Background in offensive security, AppSec, or trust & safety operations.
- Experience publishing or reproducing public adversarial-ML research.
- Multilingual fluency for cross-language attack testing.
Skills
- Adversarial prompting
- Red-team analysis
- Policy taxonomy
- Failure documentation
- Secure Code Review AI evaluation
- Cybersecurity
- AI evaluation
- Rubric writing
- Expert review
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$40 per hour
A leading tech firm is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve technical problems. This role requires at least 2 years of hands-on experience in areas such as penetration testing and incident response. Candidates...SuggestedHourly payRemote work- ...Senior Software Engineer — AI Coding Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing... ...failure modes (compile error, runtime crash, off-by-one, security issue) with severity scores. Document recurring...SuggestedRemote jobHourly payFor contractors10 hours per week
- Obsidian is looking for expert Evaluators in Finance operations/audit support to review AI-generated work products for accuracy and quality. This remote hourly position requires a minimum of 5 years in finance and fluency in English. Your role will involve evaluating outputs...SuggestedRemote jobHourly payWork at office
- Obsidian is hiring expert Evaluators in Investment analysis / valuation / credit to review AI-generated work products for accuracy and quality. This remote, hourly position requires deep subject-matter expertise and professional fluency in English to provide structured...SuggestedHourly payWork at officeRemote work
- AuraOne is seeking a Prompt Injection Security Evaluator for a remote contractor role. You will design adversarial prompts, document failures with reproducible steps, and score defenses across single-turn and multi-turn conversations. The role emphasizes rigorous reporting...SuggestedRemote jobFor contractors
$100 per hour
...This Role Actually Is You will assess how AI coding agents behave in real-world scenarios —... ...correctness. What You’ll Be Doing Evaluate AI-generated coding interactions end-to-... ...needing to fully execute or deeply review every line Comfortable giving direct...Contract workImmediate start- ...Cloud Security AI Evaluator is a remote evaluation track for reviewing cloud security ai evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team...Remote jobHourly payFor contractors10 hours per week
$150 per hour
...This Role Actually Is You will assess how AI coding agents behave in real-world scenarios —... ...correctness. What You’ll Be Doing Evaluate AI-generated coding interactions end-to-... ...needing to fully execute or deeply review every line Comfortable giving direct...Contract work$208k - $416k
...steps. Our partner is looking for an AI Interaction Evaluator based in the United States. This is... ...opportunity focused on evaluating how modern AI coding agents interact with experienced... ...: Evaluate AI interactions: Review AI-generated coding interactions end to...Hourly payContract workTemporary workImmediate startRemote work10 hours per weekFlexible hours- ...Corporate and Securities Law AI Evaluator is a remote review track for evaluating AI outputs in legal review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
- ...We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks... ...lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation...Contract work
- ...We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks... ...lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation...Contract work
- AuraOne seeks a VFX Quality Evaluator to review AI outputs across vfx review workflows. Reviewers grade craft, constraints, and tooling adherence; flag production-readiness issues; and document the corrected approach so the modeling team can train on it. This remote contractor...Remote jobFor contractors10 hours per week
- AuraOne is seeking an Indigenous Language AI Evaluator to remotely review indigenous language AI evaluation prompts and responses against our quality rubric. You will compare paired outputs, label edge cases, and provide structured feedback to help retrain the model. This...Remote jobFor contractors10 hours per week
$90 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...from those that merely look correct. Evaluate responsive behavior and semantic quality,... ...preference labeling, or model evaluation work. Code review or technical assessment background....Contract workSummer workLocal areaRemote work$60 per hour
...contribute to developing cutting-edge AI systems, while enjoying the... ...-art AI models on tasks like evaluating AI-generated quantitative... ...and well-documented analytical code. Provide feedback that directly... ..., with comfort writing and reviewing analytical code end-to-end....Hourly payFull timeRemote workFlexible hours- ...remote, hourly contractor role supporting AI data and language projects on a project... ...support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy,... ...AI products, with deep expertise in Arabic-native and secure, sovereign solutions....Hourly payFor contractorsRemote workFlexible hours
$80 - $120 per hour
...the accuracy, rigor, and domain quality of AI-generated legal work products related to... ...apply deep subject-matter expertise to review documents, spreadsheets, and slide decks,... ...for legal accuracy and domain rigor. Evaluate outputs against domain-specific quality rubrics...Hourly payWork at officeLocal areaRemote work$80 - $120 per hour
...Review AI-generated finance operations and audit support materials, including documents, spreadsheets, and slide decks, using your domain... ..., rigor, and overall quality. Key Responsibilities Evaluate AI-generated work products against domain-specific quality rubrics...Hourly payWork at officeRemote work- A leading AI company is seeking detail-oriented linguists with native speaker fluency in Vietnamese. This remote... ...working on AI-related projects such as prompt evaluation, video content understanding, and text review. Ideal candidates will have strong English skills, attention...Remote jobFreelance
- Obsidian is hiring expert Evaluators for a remote, hourly role focused on Compliance and regulatory response with financial-services AI. You'll assess AI-generated work products for accuracy and domain quality, leveraging your expertise. The ideal candidate has over 5 years...Remote jobHourly payWork at office
- AuraOne is seeking a remote Secrets Detection Security Evaluator to review model outputs, apply structured rubrics, and label edge cases. You will compare paired responses, justify the stronger answer, and calibrate your scoring against gold-standard examples. The role...Remote jobFor contractors
$60 - $70 per hour
...technical talent with leading AI research labs. Headquartered in... ...Role Responsibilities Evaluate AI-generated responses for safety... ..., and overall quality. Review content involving misinformation... ...policy, scientific research, security, or a related field. ~ Excellent...Contract workSummer workRemote work- ...About The Opportunity We are seeking detail-oriented human reviewers with a strong understanding of their local cultural context to support a range of AI training and evaluation projects . In this role, you will work across diverse task types, including evaluating...Extra incomeFull timeFor contractorsFreelanceLocal area10 hours per week
$40 per hour
A cybersecurity training company is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This is a remote role offering flexibility in project selection and scheduling. Ideal candidates will have over 2 years of cybersecurity...Hourly payRemote work- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...Work at officeRemote work
$14.5 per hour
...AI Web Search Evaluator Welo Data works with technology companies to provide datasets that are high-quality, ethically sourced, relevant,... ...volumes can vary week to week. Some weeks there is more data to review, other weeks less. Start Date: ASAP Employment Type:...Hourly payPart timeCurrently hiringImmediate startRemote workWork from home10 hours per weekFlexible hours- ...MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers...Contract workTemporary workImmediate startRemote work
$14.5 per hour
...AI Web Search Evaluator Unlock the Power of the Internet! Are you curious, tech-savvy, and passionate about improving online search experiences... ...your home! As a Web Search Evaluator, you will: Review and assess internet search results, ensuring users receive...Bi-weekly payHourly payPart timeImmediate startRemote workWork from homeFlexible hours$85 per hour
...and technical talent with leading AI research labs. Headquartered in San... ...Dorsey. Position: iOS Engineer (Coding Agent Experience) Type: Contract... ...frontier AI coding agents to complete and evaluate complex engineering tasks. ~Review model-generated mobile application...Contract workPart timeSummer workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Secure Code Review AI Evaluator [Remote]. Be the first to apply!








