Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI

New York, NY
  • Remote job

Anyone AI Labs is seeking a Research Scientist for LLM Evaluations and Benchmarking. This remote role spans LatAm/US and requires designing robust evaluation methodologies for frontier models, building benchmarks across reasoning, coding, agents, tool use, and multi-modal tasks. The candidate will lead expert pools, validate ground truth, and publish results in venues like NeurIPS, ICLR, and ACL, with strong English proficiency and preference for Spanish speakers. #J-18808-Ljbffr Anyone AI

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Remote LLM Evaluation Scientist: Benchmarking Models in New York, NY vacancy
  • $245k - $315k

     ...Applied Research Scientist, LLM Evaluation & Post-Training Innodata is expanding its GenAI research...  ...strategies, and feedback signals influence model improvement. This role is ideal for...  .... Your work may include designing benchmark datasets, developing evaluation... 
    Remote work

    Innodata Inc.

    United States
    2 days ago
  •  ...edge foundation AI models and end-to-end products...  ...us!Why this role?Evaluation is critical to...  ...infrastructure to measure LLM progress.As a Senior Research Scientist, Model Evaluation,...  ...new evaluation benchmarks that push the limits...  ...offices if you are remote, plus an annual company... 
    Remote work
    Full time
    Work at office
    Local area
    Home office

    Cohere

    New York, NY
    2 days ago
  • Cohere is seeking a Senior Research Scientist, Model Evaluation, to create ambitious evaluation benchmarks and scale evaluation...  ...measurements and to push the frontiers of LLM evaluation. The role emphasizes...  ...research and engineering, with remote-friendly policies and... 
    Remote work

    cohere

    New York, NY
    1 day ago
  • Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New...  ...infrastructure to measure LLM progress. As a Senior Research...  ...ambitious new evaluation benchmarks that push the limits of...  ...and workspace improvement Remote‑flexible, offices in Toronto... 
    Remote work
    Full time
    Work at office
    Flexible hours

    SupportFinity™

    New York, NY
    2 days ago
  • $195.2k - $262.2k

     ...enterprises from data and model training through...  ...Factory needs scientists who can turn...  ...programs in efficient LLM and VLM inference...  ...impact. Invent, evaluate, and productionize...  ...artifacts, widely used benchmarks, high-quality...  ...secondary caregivers. Remote work reimbursement... 
    Remote work
    Temporary work
    Immediate start

    Nebius

    Palo Alto, CA
    1 day ago
  • $300k

     ...is looking for a Machine Learning Scientist 4 - Generative Models, Evaluation based in United States. This is...  ...and automated evaluation. Build LLM-based evaluators: Research and develop...  ...receiving flexible time off. Remote work: Full-time remote position within... 
    Remote work
    Full time
    Flexible hours

    jobgether

    United States
    6 days ago
  •  ...computational problem solving to improve and evaluate large language models. You will design rigorous math...  .... Help define new evaluation benchmarks based on mathematics curricula spanning...  ...ability to work independently in a remote setting. Technical requirements:... 
    Remote work
    Contract work
    For contractors
    Freelance

    SaidGig

    United States
    a month ago
  •  ...States Digital Space LLC offers a remote contract role focusing on fine-tuning large language models through rigorous mathematical...  ...broken down complex problems for evaluation tasks. You will collaborate with LLM researchers on benchmarks spanning undergraduate to PhD topics... 
    Remote job
    Contract work

    United States Digital Space LLC

    New York, NY
    2 days ago
  • $100 - $120 per hour

     ...computer vision and language modeling. This role focuses on training, improving, evaluating, and deploying deep...  ...generators. LLM post-training and behavioral...  ...learning, data ordering, benchmark construction,...  ...plus. Work Terms Remote, hourly engagement.... 
    Remote work
    Hourly pay
    Flexible hours

    SaidGig

    United States
    25 days ago
  • $40 per hour

     ...is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates...  ...biochemistry. This position allows full-time or part-time remote work, with projects that pay hourly starting at $40+. Strong... 
    Remote work
    Hourly pay
    Full time
    Part time

    DataAnnotation

    United States
    1 day ago
  •  ...zone is seeking an AI Research Scientist to lead applied AI research...  ...into measurable experiments in LLM evaluation and RLHF data design. You will...  ...cross-functional teams to improve model performance, safety, and usefulness. The role is remote in the United States, with... 
    Remote job
    Hourly pay
    Flexible hours

    AIToolboard

    New York, NY
    3 days ago
  • Lynker Technologies Sea Ice Model Evaluations and Applications Scientist US-MD-Suitland Job ID: 2026-1639 Type: Full-Time # of Openings: 1 Suitland Overview...  ...: Knowledge of sea ice analysis, forecasting, remote sensing, dynamics, or thermodynamics. Experience with... 
    Remote work
    Full time
    Seasonal work
    Local area

    Lynker Technologies

    Suitland, MD
    5 days ago
  • $60 - $80 per hour

     ...mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in mathematics. Experts contribute domain expertise to model training and evaluation, create real-world...  ...research outcomes Work independently on remote projects, managing your time to meet project... 
    Remote work
    Hourly pay
    Contract work
    Immediate start

    SaidGig

    United States
    more than 2 months ago
  •  ...Overview Drive the creation and evaluation of challenging STEM problems used to fine-tune and benchmark large language models. You will design multi-step...  .... This role is fully remote, contract-based, and ideal...  ...expertise to support cutting-edge LLM research and production... 
    Remote work
    Contract work
    For contractors
    Freelance

    SaidGig

    United States
    more than 2 months ago
  • $50 per hour

     ...This role focuses on improving and evaluating large language models through advanced mathematical reasoning...  .... Contribute to new evaluation benchmarks spanning curricula from early undergraduate...  ...to collaborate effectively in a remote setting. Self-motivated, able to... 
    Remote work
    Contract work
    For contractors
    Freelance

    SaidGig

    United States
    3 days ago
  •  ...complex, multi-step ML tasks for frontier models. You will design experiments, implement...  ...fall short. This full-time, fully remote role within the United States requires a...  ...and strong Python/Git skills, with a focus on RL and LLM evaluation. #J-18808-Ljbffr Obsidian
    Remote job
    Full time
    Part time

    Obsidian

    New York, NY
    2 days ago
  •  ...researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating...  ...written conclusions. You will work remotely in the United States for...  ...teams through Cincinnatus’ extended workforce model. #J-18808-Ljbffr Obsidian
    Remote job
    Full time
    Part time

    Obsidian

    San Francisco, CA
    2 days ago
  •  ...-based solutions, and perform rigorous data analysis. You will work remotely within the United States, about 35 hours per week, collaborating with researchers to ensure consistent, high-quality benchmark design and evaluation of frontier models. #J-18808-Ljbffr Mercor
    Remote job

    Mercor

    San Francisco, CA
    2 days ago
  • $70 per hour

     ...research problems that help evaluate advanced AI systems in scientific...  ...tasks that strong frontier models do not consistently solve....  ...Contribute to a scientific computing benchmark by authoring rigorous...  ...experience. Work Terms Remote, hourly engagement. Six-week... 
    Remote work
    Hourly pay
    Part time
    Immediate start

    SaidGig

    United States
    7 days ago
  • Scientist / Senior Scientist, Structure-Based Modeling Scientist / Senior Scientist, Structure-Based...  ...Nice to have: Experience benchmarking across multiple PDB...  ...for predictive modeling; Evaluate multiple structural representations...  .... We offer: A remote-first team across the... 
    Remote work
    Full time
    Flexible hours

    Deep Origin

    New York, NY
    2 days ago
  • $30 - $50 per hour

     ...Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing AI research experiments,... 
    Remote job
    Hourly pay

    REX

    New York, NY
    2 days ago
  • $30 - $50 per hour

     ...projects. The role involves end-to-end research cycles, building and evaluating LLM systems, and collaborating on dataset development. The ideal...  ..., and strong written communication skills. This is a remote, full-time position with competitive hourly compensation ranging... 
    Remote job
    Hourly pay
    Full time

    REX

    New York, NY
    2 days ago
  • $70 - $85 per hour

     ...and geophysics challenges that evaluate whether advanced AI systems...  ...focuses on creating original benchmark problems grounded in practical...  ...problems with advanced AI models and refine them until they reach...  ...terminal environment with remote compute sandboxes. Ability... 
    Remote work
    Hourly pay

    SaidGig

    United States
    16 days ago
  • $80 - $135 per hour

     ...quality reference solutions for the CritPt benchmark (arXiv:2509.26574v3), a frontier...  ...-verified reference data used to evaluate large language model performance on frontier physics reasoning...  .... Work Terms Location: Remote. Employment type: hourly. Expected... 
    Remote work
    Hourly pay
    10 hours per week

    SaidGig

    United States
    a month ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that frontier models cannot solve. Domains—depth in at least two subdomains with a... 
    Part time
    Immediate start

    Mercor

    New York, NY
    2 days ago
  • $50 - $70 per hour

     ...About the job Remote | Applied Physicist (AI Benchmarking) - $50-$70/hour We are sharing a specialised part-...  ...in applied physics, mathematical modelling, experimental analysis, or computational...  ...-choice assessment content, evaluate scientific accuracy and solution quality... 
    Remote work
    Full time
    Part time
    For contractors
    10 hours per week
    Flexible hours

    24-MAG LLC

    New York, NY
    13 days ago
  • $100 - $120 per hour

     ...role combines end-to-end model development, from...  ...training and TRADES. Evaluating robust accuracy under standard...  ...parameter counts. LLM Post-Training and...  ...evaluation, including benchmark construction, contamination...  ...contributions. Work Terms Remote role. Hourly,... 
    Remote work
    Hourly pay
    Temporary work
    Flexible hours

    SaidGig

    United States
    3 days ago
  •  ...based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical...  ...applied experience, and strong communication skills. This role is remote, offering an opportunity to work independently in a fast-paced... 
    Remote work

    Kake

    Austin, TX
    2 days ago
  •  ...team combines deep expertise in model innovation and systems...  ...We're looking for a Research Scientist who can define what "better"...  ...'s model families, build the evaluation infrastructure to measure it...  ...meaningful progress, not just benchmark performance. Develop novel quantitative... 

    Sanas

    Palo Alto, CA
    5 days ago
  •  ...and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is perfect for someone with deep research expertise who... 

    Vals AI

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote LLM Evaluation Scientist: Benchmarking Models. Be the first to apply!