Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer, Model Evaluations

Full-time

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do.

Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on. You'll design and implement evaluations across the full spectrum of Claude's capabilities and personality, and build the infrastructure that runs them reliably at scale. You'll partner closely with researchers throughout the lifecycle of a new capability — from defining what to measure, to running the eval against live training checkpoints, to interpreting the results. The goal is to make Anthropic the leader in extremely well-characterized AI systems, with performance that is exhaustively measured and validated across the tasks that matter.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Research Engineer, Model Evaluations in Remote vacancy
  • $50 - $100 per hour

     ...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback, and code...  ...role focuses on code generation workflows and model evaluation; prior AI experience is not required. Key... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    4 days ago
  •  ...Research Engineer - Code Generation & Model Evaluation is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    2 days ago
  • $60 - $90 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 day ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign to models. Research leads on this team collaborate with engineers in one of our focus areas: CBRN, Cyber, Autonomy... 
    Suggested
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    New York, NY
    5 days ago
  • $70 - $80 per hour

     ...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports... 
    Suggested
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $100 per hour

     ...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas such as...  ...on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness...  ...NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    24 days ago
  • $20 - $36 per hour

     ...Role Overview Evaluate generative music AI across a wide range of genres, applying your knowledge of Hungarian music and lyrics to detailed quality standards. You will work in both Hungarian and English to help assess the quality, originality, and naturalness of AI-generated... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    7 days ago
  •  ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets... 
    For contractors
    Remote work

    SaidGig

    United States
    22 days ago
  •  ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine learning models and behavioral algorithms that drive our autonomous vehicle... 
    Full time
    Temporary work
    Relocation package

    Zoox

    California
    4 days ago
  • $70 - $90 per hour

     ...kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical...  ...including NKI, Pallas, or TPU. Background in compiler engineering, MLIR, or intermediate-representation lowering. Knowledge... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    24 days ago
  • $20 per hour

    SupportFinity™ is seeking an Editorial Proofreader to evaluate AI models and improve their quality through expert writing and editing skills. This role can be part‑time or full‑time, allowing for a flexible schedule and project selection. Applicants must be fluent in English... 
    Remote job
    Hourly pay
    Full time
    Part time
    Flexible hours

    SupportFinity™

    Sioux Falls, SD
    2 days ago
  • $20 per hour

    SupportFinity™ is looking for an Editorial Proofreader to join our team for AI model training. In this remote role, you'll evaluate AI chatbots and enhance model quality. Candidates should have fluency in English and strong editing skills. This position can be full‑time... 
    Remote job
    Hourly pay
    Full time
    Contract work
    Part time

    SupportFinity™

    New York, NY
    6 days ago
  •  ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    6 days ago
  • $60 per hour

     ...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    6 days ago
  • $100 - $150 per hour

     ...considered for future projects evaluating how well AI systems perform...  ...human-produced analyses and models, document decisions in...  ...A/B test write-ups, feature engineering, and technical reports or notebooks...  ...at a leading technology, research, or quantitative firm, such... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  •  ...computational problem solving to improve and evaluate large language models. You will design rigorous math...  ...customers Accelerate frontier AI research by contributing high quality data...  ...mathematics at the level expected for engineering entrance exams and for graduate or... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  •  ...happens: every product built on a model is bounded by what it costs...  ...: ~ This role builds the evaluation and decision systems that...  ...token consumption. You will research and prototype adaptive model-...  ...closely with the MaaS and platform engineering teams on production... 
    Full time

    Bitdeer Technologies Group

    Austin, TX
    18 days ago
  • $60 - $80 per hour

     ...application, not a posting for one specific role. Role Overview Qualified experts may support AI research by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities Train and evaluate AI models in... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $110 per hour

     ...your background and interests. Role Overview Qualified physicians may contribute to AI research projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models in medicine. Create tasks... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $60 - $90 per hour

     ...Overview Apply hands-on mechanical engineering judgment to improve how advanced AI models reason through real-world...  ...work. You will partner with AI research and program management teams to...  ...solutions, and create rigorous evaluations grounded in industry practice.... 
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    18 hours ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    28 days ago
  • $65 - $105 per hour

     ...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world software development...  ...You will work closely with an AI research and program management team to...  ...define high-quality engineering work, evaluate model performance, and turn expert... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    28 days ago
  • $90 - $175 per hour

     ...software quality assurance expertise to evaluate technical AI outputs and help improve how...  ...Qualifications Professional experience as a QA Engineer, SDET, Test Engineer, QA Analyst, or in a...  ..., labeling, RLHF, AI response or model evaluation, or rubric-based grading. Software... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    10 days ago
  • $60 - $90 per hour

     ...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong from... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    18 days ago
  • $70 - $110 per hour

     ...this senior clinical medicine role, you will partner with an AI research team to evaluate medical knowledge tasks, define high-quality clinical standards, and develop benchmarks that measure meaningful model improvement. Key Responsibilities Review clinical... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    28 days ago
  •  ...Apply your dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical, remote opportunity with no direct patient care, focused on ensuring AI-generated medical outputs reflect real... 
    Hourly pay
    Remote work

    SaidGig

    United States
    9 days ago
  •  ...Help establish the clinical reference standards used to evaluate advanced AI systems on real radiology studies. This role draws on the...  ...Contribute challenging cases that test the limits of strong AI models. Qualifications Board-certified or board-eligible radiologist... 
    Hourly pay
    Remote work

    SaidGig

    United States
    4 days ago
  • General Motors, through Embodied AI, seeks a Senior Engineer to measure and visualize AV model performance. You will design and implement evaluation workflows, collaborate across Data, Infra and validation teams, and connect offline evaluation with real-world behavior.... 
    Remote job

    General Motors

    Austin, TX
    5 days ago
  • $50 - $70 per hour

     ...Help improve frontier AI systems by evaluating the quality of professional work products across documents, presentations, spreadsheets, and other written materials. Your careful assessments and written feedback will help shape how AI produces accurate, clear, and complete... 
    Hourly pay
    Remote work

    SaidGig

    United Kingdom
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer, Model Evaluations. Be the first to apply!