Prompt Injection Security Evaluator [Remote]
AuraOne Human Data
- Remote job
Prompt Injection Security Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team can patch the gap.
Why this role matters
Adversarial evaluation is how AuraOne hardens AI models before they ship to customers. Reviewers think like attackers and write up failures with enough rigor that the modeling team can reproduce, fix, and regress-test them.
Responsibilities
- Design adversarial prompts that probe known weakness classes (jailbreak, policy bypass, prompt injection) for Prompt Injection Security Evaluator assignments.
- Document every successful attack with reproduction steps and the policy clause it violated.
- Score model defenses across single-turn and multi-turn conversations.
- Triage emerging attack vectors and route them to the safety team with severity ratings.
- Maintain a personal library of attack patterns and propose new red-team rubrics.
- Calibrate against the broader red-team cohort to keep coverage and severity consistent.
Qualifications
- Demonstrated experience red-teaming AI systems, security research, or adversarial ML work for Prompt Injection Security Evaluator work.
- Strong written communication — your reports become the patch ticket.
- Comfort working in policy-grey areas with clear documentation of what was attempted and why.
- Familiarity with prompt-injection, jailbreak, and policy-bypass taxonomies.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Construct a 5-turn adversarial conversation that bypasses a specific policy clause and write up the patch ticket.
- Score a model's defenses against a known jailbreak pattern across 20 variants.
- Propose a new red-team rubric category after spotting an emerging attack vector.
- Reproduce a failure another reviewer reported and confirm the severity tag.
Nice to have
- Background in offensive security, AppSec, or trust & safety operations.
- Experience publishing or reproducing public adversarial-ML research.
- Multilingual fluency for cross-language attack testing.
Skills
- Adversarial prompting
- Red-team analysis
- Policy taxonomy
- Failure documentation
- Prompt Injection Security evaluation
- Security review
- Adversarial testing
- Threat modeling
- Prompt
- Injection
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$35 - $40 per hour
A leading evaluation organization seeks a Generalist Evaluator Expert for a remote, flexible contract role. This position involves designing prompts for language model evaluations, defining standards, and conducting assessments. Ideal candidates should have strong writing...SuggestedRemote jobContract workFlexible hours- ...Browser Extension Security Security Evaluator is a remote evaluation track for reviewing browser extension security security evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...SuggestedRemote jobHourly payFor contractors10 hours per week
$60 - $80 per hour
...Operations Research Model Prompt Evaluator is a remote review track for evaluating AI outputs across operations research model prompt research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and...SuggestedRemote jobFor contractors10 hours per week- Feitong Buke is hiring a Lead AI Trainer to oversee and enhance the quality of AI model dialogues with users. The role involves reviewing datasets for accuracy, providing feedback to annotators, and validating AI model outputs to ensure high production quality. Candidates...SuggestedRemote jobFull time
$30 - $45 per hour
...Job Description: Web Browsing Evaluator remote $30 - $45/hour pay... ...Navigate web pages according to detailed prompts and project specifications, ensuring browser... ...and adhere to all project compliance and security protocols. Preferred Qualifications...SuggestedFor contractorsRemote work- ...linguists with native Turkish fluency for a remote freelance opportunity. This role will focus on AI-related projects that involve prompt evaluation and multimedia content understanding. Ideal candidates will have a strong command of English and experience as reviewers or...FreelanceRemote work
$100 per hour
...looking for a highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as... ...tools like Cursor or similar AI-first IDEs ~Prior exposure to prompt design or evaluation workflows ~Experience mentoring senior...Contract workImmediate start$25 - $35 per hour
...Professional Domain Expert (Real Estate and Leasing) to work part-time and remotely. Responsibilities include writing prompts for real estate scenarios, evaluating AI outputs, and providing evidence-based evaluations. The ideal candidate has expertise in real estate or...Remote jobHourly payPart time- Freelance Luxury Brand Evaluator Automotive Project - Greater Los Angeles Are you a luxury automobile enthusiast who appreciates the finer... ...unbiased, honest feedback without personal biases and be prompt in filling out online surveys. Benefits: This is a freelance, project...FreelanceWorldwideFlexible hours
- Rex.zone is seeking a Senior AI data annotator to perform data labeling and evaluation for NLP tasks, RLHF assessments, and prompt QA to improve training data quality and model performance. This is a US-based remote, full-time role aligned with Miami talent demand. You...Remote jobFull time
- ...fueled by our partnership with EQT. Website: Linkedin Job Title: Evaluator - Political Science Location: Remote (USA) Job Type: Contract... ...is the usefulness in addressing the question posed in the user prompt. The goal is to assess quality, safety, and utility—...Contract workRemote workWorldwide
- Dorado is seeking a Speech AI Evaluation Specialist to support the improvement of AI-generated content in Vietnamese (USA). This is a freelance... ...in short voice conversations with AI models, follow prompts, evaluate responses for relevance, accuracy, and clarity, and...Remote jobPart timeFreelanceImmediate startFlexible hours
- Dorado seeks Speech AI Evaluation Specialists based in Malaysia for remote, part-time freelance work. You will engage in brief voice conversations with AI models, follow scripted prompts, and rate responses on relevance and quality. Strong Chinese Simplified and good English...Remote jobPart timeFreelanceImmediate startFlexible hours
- Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in Italian. This freelance, remote role offers... ...will conduct short voice conversations with AI models, follow prompts, and evaluate responses for relevance, accuracy and clarity, providing...Remote jobPart timeFreelanceImmediate startWork from homeFlexible hours
$190 per hour
...hourLocation:RemoteCommitment:20+ hours/week Role Responsibilities Evaluate AI systems on complex personal workflows, including personal health, travel, and activity planning. Design realistic prompts for complex personal-life tasks using Google Drive, Expedia, Notion...Remote jobHourly payFor contractorsSummer workTrial period$8 per hour
Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in Thai or Chinese Simplified. This freelance,... ...will engage in short voice conversations with AI models, follow prompts, and rate AI responses for quality and accuracy. Ideal...Remote jobPart timeFreelanceImmediate start$18 per hour
Dorado is seeking a Speech AI Evaluation Specialist to help improve AI-generated content in German. This freelance, part-time role offers... ...participate in short voice conversations, follow scenario prompts, and evaluate AI responses for relevance, accuracy, and clarity...Remote jobHourly payPart timeFreelanceImmediate startWork from home10 hours per weekFlexible hours- ...States English-language work 20+ hours per week About AI Response Evaluation AI training is the human side of building artificial... ...Sufficient personal application history to create and assess retrieval prompts is helpful Who Should Apply This role may suit people who are...Part timeFor contractorsRemote work
$5,083 per month
...11, 2026. This position specializes in evaluation functions under the general supervision... ...and Information Respond to all inquiries promptly from staff, students, faculty and alumni... ...and federal crime awareness and campus security legislation, including The Jeanne Clery...Work at officeRemote work- ...a great match for you. Become a Luxury Automotive Experience Evaluator As a Luxury Automotive Experience Evaluator, you'll be invited... ...provide unbiased, honest feedback without personal biases and be prompt in filling out online surveys Benefits This is a freelance, project...FreelanceFlexible hours
- ...Policy Compliance Risk Evaluator is a remote red-team track for stress... ...systems against adversarial prompts. Reviewers craft attack... ...jailbreak, policy bypass, prompt injection) for Policy Compliance Risk Evaluator... ...red-teaming AI systems, security research, or adversarial ML...Remote jobHourly payFor contractors10 hours per week
- ...Security Policy Bypass Security Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
$42 per hour
...months Commitment: 20+ hours/week Role Responsibilities Evaluate AI-generated music across various genres and rate it against... ...similarities. Rate lyrics based on quality, creativity, prompt adherence, and originality. Assess whether lyrics sound natural...Remote jobContract workSummer workImmediate start$42 per hour
...months Commitment: 20+ hours/week Role Responsibilities Evaluate AI-generated music across various genres for quality and... ...identify similarities. Rate lyrics based on quality, creativity, prompt adherence, and originality. Assess naturalness of lyrics,...Remote jobContract workSummer workImmediate start$70 per hour
...Source material from published papers, Kaggle datasets, open-source repositories, or designed scenarios. Write scientific prompts based on sourced input to challenge AI models . Build grading criteria to define correct answers for scientific tasks. Calibrate...Contract workSummer workImmediate startRemote work- ...improving model performance. You will review technical outputs, craft prompts, and provide structured feedback. Ideal candidates have 3+... ...strong writing skills, and experience with data annotation or evaluation rubrics. This contractor, part-time role supports distributed...Remote jobPart timeFor contractors
- ...experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will... ...judgments, and data integrity, guiding model refinements and prompts for high-quality outputs. The role is remote, project-based, and...Remote job
- CNTXT AI is seeking a remote contractor to evaluate AI-generated financial content and develop test cases that probe analytical reasoning... ...-referencing data from balance sheets and reports; creating prompts and Excel-based evaluation tasks; and #J-18808-Ljbffr CNTXT AIRemote jobFor contractors
- ...The Mental Health Evaluator (MHE) is a critical role dedicated to delivering comprehensive, clinically appropriate, culturally competent, and trauma-informed patient behavioral health interventions in the hospital Emergency Department. Responsibilities include conducting...
- ...Content generation: writing high-quality prompts and model responses, or recording high-... ...to support AI training datasets. LLM evaluation: reviewing AI-generated responses for... ...proprietary AI products, with deep expertise in Arabic-native and secure, sovereign solutions....Hourly payFor contractorsRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Prompt Injection Security Evaluator [Remote]. Be the first to apply!





