LLM Red-Teamer for AI Model Evaluation [Remote]
$40 - $65 per hourSaidGig
- Remote job
Join a high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this role, you will leverage your expertise to train next-generation AI systems, shaping how models learn, reason, and perform through high-quality, real-world input. Key Responsibilities
- Develop complex, adversarial multi-turn conversations and task-based scenarios aligned with detailed project specifications.
- Author clear, precise evaluation rubrics to rigorously assess model responses against defined behavioral targets.
- Iteratively test conversations and tasks against frontier LLMs, escalating difficulty and nuance until the desired quality threshold is achieved.
- Deliver comprehensive task packages, including transcripts, target behaviors, binary rubrics, and supporting rationale or evidence.
- Validate LLM outputs, documenting model strengths and failure modes relative to the project specification.
- Maintain calibration with team leads and quality control contacts as project requirements evolve.
- Contribute independently, producing high-quality deliverables at a steady and consistent pace.
- Exceptional written English skills, with clarity, precision, and strong structural organization.
- Prior experience in AI human data environments (RLHF, SFT, evaluations, annotation, or prompt engineering) is preferred.
- Deep familiarity with large language models, including the ability to anticipate and identify common failure patterns.
- Demonstrated ability to work autonomously, interpreting and executing complex specifications with minimal oversight.
- Proven critical thinking and meticulous attention to detail.
- Experience designing evaluation items or rubrics is advantageous.
- Background in writing-intensive or analysis-centric fields such as research, editorial, technical writing, or quality assurance is a plus.
This is a contractor position with remote work flexibility. Experts are expected to submit a minimum number of tasks per week.
CompensationCompensation is output-based, ranging from $40 to $65 per hour, depending on the expert''s experience and workflow.
EligibilityWe typically fill roles within 48 hours and are looking for experts ready to start immediately. Selected candidates are expected to begin their first tasks within 24, 48 hours of completing onboarding.
$40 - $65 per hour
...Join a high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this role, you will leverage your expertise to train next-generation AI systems, shaping how models learn, reason, and perform through high-quality...SuggestedRemote jobHourly payFor contractorsImmediate start$60 - $90 per hour
...where frontier language models appear competent but quietly... ...findings into robust evaluation benchmarks. Key... ...engineering, security, or AI evaluation. Proven ability... ...or ML systems, through red teaming, adversarial... ...Strong familiarity with LLM capabilities, limitations...SuggestedHourly payFull timeFreelanceRemote work$350k
...role in shaping the future of AI-powered legal reasoning. This... ...intersection of large language models, agentic systems, and legal workflows... ...the development of rigorous evaluation frameworks to measure and... ...Advanced degree in Law (JD, LLM, SJD, PhD in Law, or equivalent...SuggestedRemote jobFull time$105 per hour
...leverage their expertise to contribute to AI research projects focused on high-... ...community of experts to refine and evaluate the capabilities of Large Language Models (LLMs) in creating impactful... ...domain-specific prompts and evaluate LLM responses for business contexts....SuggestedWork experience placementRemote workFlexible hours$60 - $150 per hour
...Role Overview Provide legal subject-matter expertise to improve and evaluate AI systems, by designing realistic legal tasks, reviewing model outputs, and giving domain-specific feedback that advances frontier AI research. This is an open application to join a Law Expert...SuggestedHourly payContract workImmediate startRemote work$80 - $110 per hour
...Overview Work on the forefront of generative AI by designing and executing real-world... ...and reasoning gaps in advanced models. You will author tasks, produce reference... ...and executable tests where applicable, run evaluations against a target model, and analyze failures...Hourly payPart timeFreelanceRemote work$85 per hour
...Role Overview GIS Analysts apply hands-on expertise in environmental assessment, GIS analysis, and renewable energy siting to evaluate AI-generated geospatial outputs and develop expert training data that improves AI understanding of environmental workflows and mapping...Hourly payContract workPart timeWork at officeRemote workFlexible hours$85 per hour
...Role Overview REC Traders evaluate AI-generated content using their renewable energy and REC trading expertise, creating expert training... ...clear, detailed feedback on AI-generated responses to help refine model behavior and domain correctness. Work independently and...Hourly payFull timeContract workPart timeFor contractorsRemote workFlexible hours$75 per hour
...Managers apply archival, library, and collections expertise to evaluate and improve AI-generated content related to records, archives, and... ...will create prompts that reflect real workplace tasks, review model outputs for accuracy and relevance, and provide clear, structured...Part timeRemote workFlexible hours- ...Physicians apply clinical judgment and frontline medical experience to evaluate AI-generated medical content, ensuring clinical accuracy, sound... ...planning. Assess clarity, relevance, and safety of model outputs in realistic care scenarios. Provide detailed, constructive...Full timeFor contractorsPrivate practiceRemote workFlexible hours
$80 - $150 per hour
...in a high-impact project that shapes the future of AI systems by applying your physics expertise. As a... ...Junior Professor), you will play a critical role in evaluating and enhancing the training of next-generation AI models, ensuring they learn and reason effectively...Remote jobHourly payFor contractors$75 per hour
...Role Overview Physics experts apply advanced physics training to evaluate AI-generated scientific content and provide detailed feedback that improves AI physical reasoning, mathematical modeling, theoretical analysis, and quantitative problem solving. This is a project...Hourly payFull timeContract workPart timeRemote workFlexible hours$220k
...This role focuses on advancing the evaluation and development of cutting-edge coding agents. You will operate at the intersection of AI research, software engineering, and model evaluation, designing the benchmarks, methodologies, and data systems that shape how next-...Full timeRemote work$90 - $130 per hour
...seasoned funds-focused legal expertise to improve how advanced AI systems read, evaluate, and negotiate investment-related contracts. In this part-... ...and give precise legal feedback that trains and refines AI models for fund formation and fund management workflows. Key...Hourly payContract workPart timeFor contractorsWork at officeRemote work$80 - $105 per hour
...Role Overview Help define how advanced AI understands and evaluates fund-related contracts by applying hands-on funds law experience to contract redlining, simulated negotiations, and model evaluation. This part-time, contractor role supports the development of AI systems...Hourly payContract workPart timeFor contractorsRemote work$80 - $135 per hour
...end reference solutions for the CritPt benchmark (arXiv:2509.26574v3). This role produces definitive solutions used to evaluate large language models on frontier physics reasoning, by solving research-level problems, auditing peer submissions, or adjudicating between competing...Hourly payRemote work10 hours per week$60 per hour
...and contribute to developing cutting-edge AI systems, while enjoying the flexibility... ...professionals to help advance AI development. AI models are increasingly capable of performing... ...state-of-the-art AI models on tasks like evaluating AI-generated quantitative analysis,...Hourly payFull timeRemote workFlexible hours$40 per hour
...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates...Hourly payRemote workFlexible hours$40 per hour
...A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative analysis and provide impactful feedback. This fully remote role allows for flexible scheduling and competitive pay starting at $40 per hour. Candidates...Hourly payRemote workFlexible hours$40 per hour
A leading AI development firm in Michigan is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the evolution of AI models. Candidates should have a background in data science, statistics, or similar fields, with at least...Hourly payRemote work$40 per hour
...A leading AI development firm is seeking experienced quantitative professionals to join their team remotely. The role involves evaluating AI-generated quantitative work, providing insights, and shaping the future of AI systems. Candidates should have over two years of...Hourly payRemote work$40 per hour
...data annotation company is seeking professionals in quantitative fields to enhance AI development. This fully remote role allows individuals to set flexible schedules while evaluating AI-generated analyses and solving complex quantitative problems. Candidates should have...Hourly payRemote workFlexible hours$40 per hour
...A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background...Hourly payRemote workFlexible hours$40 per hour
...A forward-thinking AI solutions company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting-edge... ...skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr...Hourly payRemote workFlexible hours$40 per hour
A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects...Hourly payRemote work$40 per hour
...A leading AI company in the United States is seeking experienced quantitative professionals to evaluate and validate AI-generated analytical work. This fully remote position allows you to set your own schedule, with competitive hourly pay starting at $40 USD. Responsibilities...Hourly payRemote work$40 per hour
A data science team is seeking experienced quantitative professionals to evaluate AI-generated work and contribute to the development of cutting-edge AI systems. This fully remote position offers flexible scheduling and competitive hourly pay starting at $40+. Ideal candidates...Hourly payRemote workFlexible hours$40 per hour
...An innovative AI development company is seeking experienced quantitative professionals to contribute to AI advancements. This fully remote role involves evaluating AI-generated analyses and ensuring they are technically accurate and valid in real-world scenarios. Candidates...Hourly payRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting...Hourly payRemote work$40 per hour
A leading AI development company is seeking experienced quantitative professionals for a remote role. Candidates will evaluate AI-generated quantitative work, solve complex problems, and provide valuable feedback. The ideal candidate has 2+ years of experience in a quantitative...Hourly payFull timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Red-Teamer for AI Model Evaluation [Remote]. Be the first to apply!
- junior javascript developer remote Remote
- business development manager remote Remote
- help desk remote Remote
- remote nurse auditor Remote
- loan officer remote Remote
- senior accountant remote Remote
- remote linux administrator Remote
- remote utilization review nurse part time Remote
- remote christian Remote
- remote healthcare recruiter Remote


