Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer, Model Evaluations

SignalAI

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About The Role We\'re looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do. Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on. You\'ll design and implement evaluations across the full spectrum of Claude\'s capabilities and personality, and build the infrastructure that runs them reliably at scale. You\'ll partner closely with researchers throughout the lifecycle of a new capability — from defining what to measure, to running the eval against live training checkpoints, to interpreting the results. The goal is to make Anthropic the leader in extremely well-characterized AI systems, with performance that is exhaustively measured and validated across the tasks that matter. Key responsibilities Design and run new evaluations of Claude\'s capabilities — reasoning, agentic behavior, knowledge, safety properties — and produce visualizations that make the results legible to researchers and decision-makers Build and harden the distributed eval execution platform so hundreds of evals run reliably against checkpoints throughout production RL training runs Own the dashboards researchers and leadership use to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss Debug anomalous eval results mid-training-run, determine whether the cause is a model change or an infrastructure issue, and communicate the answer clearly under time pressure Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations Partner with research teams across the full lifecycle of a new capability — from defining what to measure to interpreting results as training progresses Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks Communicate evaluations and their results to internal stakeholders and, where appropriate, external audiences Minimum Qualifications Strong Python programming skills, including production or research infrastructure Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale Clear written and verbal communication, especially when explaining technical results to non-specialists Comfort operating in an on-call or production-support capacity when training runs are live Care about the societal impacts of your work and an interest in steering powerful AI to be safe and beneficial Preferred Qualifications Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding Background in data visualization and a track record of building dashboards people actually trust and use Experience developing robust evaluation metrics for language models Experience with observability, monitoring, or experiment-tracking systems Background in statistics and experimental design Experience with large-scale dataset sourcing, curation, and processing Experience running or supporting ML training infrastructure A bias toward picking up slack and operating flexibly across team boundaries Enjoy pair programming — we love to pair Representative projects Stand up a new eval that tests a specific reasoning capability from scratch — define the task, build the dataset, implement the scoring, validate against known signals, and ship a dashboard that makes the result legible Diagnose a mid-training regression: an eval suite returns anomalous numbers, and you need to determine within hours whether it\'s the model, the harness, the data, or the infrastructure Take a flaky distributed eval pipeline and make it boring — better retries, better observability, faster feedback to researchers Partner with a research team on a new capability area, helping them articulate what "good" looks like and translating that into measurable artifacts The annual compensation range for this role is listed below. For sales roles, the range provided is the role\u2019s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary

$500,000 - $850,000 USD

Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren\'t able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you\'re interested in this work. We think AI systems like the ones we\'re building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you\'re ever unsure about a communication, don\'t click any links—visit anthropic.com/careers directly for confirmed position openings. How We\'re Different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We\'re an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates\' AI Usage: Learn about our policy for using AI in our application process. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. How do you pronounce your name? We believe that AI will have a transformative impact on the world, and we’re seeking exceptional candidates who collaborate thoughtfully with Claude to realize this vision. At the same time, we want to understand your unique skills, expertise, and perspective through our hiring process. We invite you to review our AI partnership guidelines for candidates and confirm your understanding by selecting “Yes.” Why do you want to work at Anthropic? (We value this response highly - great answers are often 200-400 words.) Add a cover letter or anything else you want to share. Please ensure to provide either your LinkedIn profile or Resume, we require at least one of the two. #J-18808-Ljbffr SignalAI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Research Engineer, Model Evaluations in New York, NY vacancy
  • Anthropic in the United States is hiring Research Engineers to build evaluations that quantify Claude's capabilities and measure reasoning, knowledge, and safety properties at scale. You will design end-to-end experiments, define metrics, and develop the infrastructure... 
    Suggested

    SignalAI

    New York, NY
    2 days ago
  •  .... We build cutting‑edge foundation AI models and end‑to‑end products that are designed...  ...our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate...  ...and Paris. Join us! Why this role? Evaluation is critical to making progress in... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Home office

    cohere

    New York, NY
    1 day ago
  •  .... We build cutting-edge foundation AI models and end-to-end products that are designed...  ...our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate...  ...and Paris. Join us!Why this role?Evaluation is critical to making progress in... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • Runway in New York is hiring a Research Engineer to lead robotics initiatives linked to world model development. The position entails full-stack engagement in robot learning, from data collection to physical evaluation of learned policies. Applicants should have hands-on... 
    Suggested

    runwayml.com

    New York, NY
    1 day ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign to models. Research leads on this team collaborate with engineers in one of our focus areas: CBRN, Cyber, Autonomy... 
    Suggested
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    New York, NY
    20 hours ago
  •  ...enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world...  ...the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft... 
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • A leading AI development company is seeking experienced quantitative professionals for remote work evaluating AI-generated quantitative analysis. Ideal candidates will have a robust background in fields like data science, economics, or biostatistics, with at least 2 years... 
    Remote work

    DataAnnotation

    New York, NY
    4 days ago
  • $40 per hour

    A growing technology company is seeking an R&D Biologist to join their team. This role involves training AI models by evaluating chatbot outputs on complex biology questions. Applicants should possess strong expertise in biology and related fields. You will work remotely... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    4 days ago
  • $40 per hour

     ...data annotation company is seeking a Statistician to join their team. This remote role involves training AI models by posing complex mathematical problems, evaluating outputs, and assessing the model's performance. Candidates must be detail-oriented and proficient in... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    4 days ago
  • We are seeking an expert to evaluate and improve our AI models through comprehensive testing and analysis. You will be responsible for designing evaluation...  ...for bias, fairness, and accuracy Collaborate with ML engineers to implement improvements Document findings and... 

    MERIT Beauty

    New York, NY
    2 days ago
  • Preference Dataset QA Reward Model Evaluator is a remote evaluation track for reviewing preference dataset qa reward model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured... 
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AuraOne

    New York, NY
    3 days ago
  • AuraOne is seeking a remote Rubric Calibration Model Evaluation Specialist contractor in the United States. The role focuses on evaluating rubric calibration model outputs, labeling issues, and providing structured feedback to retrain the model. You will compare paired... 
    Remote job
    For contractors

    AuraOne

    New York, NY
    1 day ago
  • Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts... 
    Remote job

    Dorado

    New York, NY
    19 hours ago
  • $20 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San Francisco,...  ...tools. Generate high-quality human evaluation data by identifying response strengths,...  ..., and completeness of responses. Ensure model responses align with expected conversational... 
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    5 days ago
  • AuraOne is seeking a remote Preference Dataset QA Reward Model Evaluator to assess model outputs against a versioned rubric and provide structured feedback. You will compare paired responses, label issues, and document edge cases for retraining the model. As an independent... 
    Remote job
    For contractors

    AuraOne

    New York, NY
    3 days ago
  • Productive Playhouse seeks AI Evaluators to support evaluating AI chatbots by interacting with models, assessing capabilities, safety, and usefulness. This is a project-based, task-based engagement with flexible hours and batch deliveries. Open to freelancers outside the... 
    Remote job
    Freelance
    Flexible hours

    Triwill Group

    New York, NY
    3 days ago
  • AuraOne is seeking a remote Spoken Instruction Conversation Evaluator to review prompts and responses against its quality rubric. You will...  ...edge cases, and provide structured feedback for retraining the model. As an independent contractor, you will evaluate model outputs,... 
    Remote job
    Hourly pay
    Contract work
    For contractors

    AuraOne

    New York, NY
    1 day ago
  • Cohere is seeking a Senior Research Engineer, Model Evaluation, to create next‑generation evaluation methods and scalable infrastructure. You will develop benchmarks, datasets, and environments to measure frontier model capabilities, and you will push the state‑of‑the‑art... 

    cohere

    New York, NY
    4 days ago
  •  ...opportunities through their expert network. Qualified candidates should hold a MS or PhD in a relevant field and have experience in evaluating complex biology content. Strong communication skills and proficient English are essential for success in this role. #J-18808-... 
    Immediate start

    SME Careers

    New York, NY
    1 day ago
  • Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny.... 
    Remote job

    Dorado

    New York, NY
    2 days ago
  • Mercor is seeking experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. In this role, you will assess...  ...Ideal candidates have 3+ years as a music producer/audio engineer, a college degree in music, and native or near-native... 

    Mercor

    New York, NY
    5 days ago
  • AuraOne is seeking a remote Pairwise Preference Reward Model Evaluator to review prompts and responses against our quality rubric. You will compare paired outputs, label edge cases, and provide structured feedback to help retrain the modeling team. This independent contractor... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne

    New York, NY
    5 days ago
  •  ...seeking Insurance domain SMEs to join a cutting-edge AI training program. You will evaluate AI model outputs against real underwriting practice and rubrics, guiding research and engineering teams to close knowledge gaps in underwriting, claims, and risk assessment. The... 

    Obsidian

    New York, NY
    2 days ago
  • AuraOne is seeking a Formal Specification Model Evaluator to remotely review prompts and responses against our quality rubric. This is a contractor role open to US-based, with flexible hours. You will compare outputs, label edge cases, and provide structured feedback to... 
    Remote job
    For contractors
    Flexible hours

    AuraOne

    New York, NY
    4 days ago
  • SME Careers is seeking a remote Kotlin Engineer to review AI-generated responses and create...  ...optimizing AI performance, and ensuring model accuracy. The ideal candidate has a...  ...an expert network and requires critical evaluation of technical concepts. #J-18808-Ljbffr... 
    Remote job

    SME Careers

    New York, NY
    1 day ago
  •  ...are seeking a disciplined and creative Research Engineer to join an elite systematic trading group...  ...of sophisticated financial modeling and cutting-edge distributed computing....  ...Technological Advancement: Continuously evaluate emerging technologies to enhance the team... 
    Immediate start

    Objective Paradigm

    New York, NY
    5 hours ago
  • $174k - $252k

    Apply research ideas to high-impact problems by prototyping, curating datasets...  ...solving ambiguous problems.Train, evaluate, and iterate on deep neural models and reinforcement learning...  ...achieve research objectives. Influence engineering best practices by championing code... 
    Full time
    Work at office

    Google

    New York, NY
    4 days ago
  • $165k - $260k

    Senior NLP Research Engineer - Artificial Intelligence Location New York Business Area...  ...boosted decision trees, large language models, and dense vector databases. We are expanding...  ...codeDesign, train, experiment, and evaluate NLP models, algorithms and... 
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    2 days ago
  • $175k - $225k

     ...with traders and quantitative researchers to implement, refine and deploy alpha signals, evaluate and maintain trading tools, improve...  ....About the RoleAs a Research Engineer, you will be an integral member...  ...problem spaces, extract domain models, and build ergonomic,... 
    Temporary work
    Flexible hours

    DRW

    New York, NY
    2 days ago
  • $110.7k - $379.2k

    Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans rely on a healthcare...  ...our post-training team, you will design, train, evaluate, and align the models that reason about healthcare... 
    Local area
    Visa sponsorship

    Deloitte

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer, Model Evaluations. Be the first to apply!