Remote LLM Evaluation Scientist: Benchmarking Models
Anyone AI
- Remote job
Anyone AI Labs is seeking a Research Scientist for LLM Evaluations and Benchmarking. This remote role spans LatAm/US and requires designing robust evaluation methodologies for frontier models, building benchmarks across reasoning, coding, agents, tool use, and multi-modal tasks. The candidate will lead expert pools, validate ground truth, and publish results in venues like NeurIPS, ICLR, and ACL, with strong English proficiency and preference for Spanish speakers. #J-18808-Ljbffr Anyone AI
$245k - $315k
...Applied Research Scientist, LLM Evaluation & Post-Training Innodata is expanding its GenAI research... ...strategies, and feedback signals influence model improvement. This role is ideal for... .... Your work may include designing benchmark datasets, developing evaluation...Remote work- ...edge foundation AI models and end-to-end products... ...us!Why this role?Evaluation is critical to... ...infrastructure to measure LLM progress.As a Senior Research Scientist, Model Evaluation,... ...new evaluation benchmarks that push the limits... ...offices if you are remote, plus an annual company...Remote workFull timeWork at officeLocal areaHome office
- Cohere is seeking a Senior Research Scientist, Model Evaluation, to create ambitious evaluation benchmarks and scale evaluation... ...measurements and to push the frontiers of LLM evaluation. The role emphasizes... ...research and engineering, with remote-friendly policies and...Remote work
- Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New... ...infrastructure to measure LLM progress. As a Senior Research... ...ambitious new evaluation benchmarks that push the limits of... ...and workspace improvement Remote‑flexible, offices in Toronto...Remote workFull timeWork at officeFlexible hours
$195.2k - $262.2k
...enterprises from data and model training through... ...Factory needs scientists who can turn... ...programs in efficient LLM and VLM inference... ...impact. Invent, evaluate, and productionize... ...artifacts, widely used benchmarks, high-quality... ...secondary caregivers. Remote work reimbursement...Remote workTemporary workImmediate start$300k
...is looking for a Machine Learning Scientist 4 - Generative Models, Evaluation based in United States. This is... ...and automated evaluation. Build LLM-based evaluators: Research and develop... ...receiving flexible time off. Remote work: Full-time remote position within...Remote workFull timeFlexible hours- ...computational problem solving to improve and evaluate large language models. You will design rigorous math... .... Help define new evaluation benchmarks based on mathematics curricula spanning... ...ability to work independently in a remote setting. Technical requirements:...Remote workContract workFor contractorsFreelance
- ...States Digital Space LLC offers a remote contract role focusing on fine-tuning large language models through rigorous mathematical... ...broken down complex problems for evaluation tasks. You will collaborate with LLM researchers on benchmarks spanning undergraduate to PhD topics...Remote jobContract work
$100 - $120 per hour
...computer vision and language modeling. This role focuses on training, improving, evaluating, and deploying deep... ...generators. LLM post-training and behavioral... ...learning, data ordering, benchmark construction,... ...plus. Work Terms Remote, hourly engagement....Remote workHourly payFlexible hours$40 per hour
...is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates... ...biochemistry. This position allows full-time or part-time remote work, with projects that pay hourly starting at $40+. Strong...Remote workHourly payFull timePart time- ...zone is seeking an AI Research Scientist to lead applied AI research... ...into measurable experiments in LLM evaluation and RLHF data design. You will... ...cross-functional teams to improve model performance, safety, and usefulness. The role is remote in the United States, with...Remote jobHourly payFlexible hours
- Lynker Technologies Sea Ice Model Evaluations and Applications Scientist US-MD-Suitland Job ID: 2026-1639 Type: Full-Time # of Openings: 1 Suitland Overview... ...: Knowledge of sea ice analysis, forecasting, remote sensing, dynamics, or thermodynamics. Experience with...Remote workFull timeSeasonal workLocal area
$60 - $80 per hour
...mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in mathematics. Experts contribute domain expertise to model training and evaluation, create real-world... ...research outcomes Work independently on remote projects, managing your time to meet project...Remote workHourly payContract workImmediate start- ...Overview Drive the creation and evaluation of challenging STEM problems used to fine-tune and benchmark large language models. You will design multi-step... .... This role is fully remote, contract-based, and ideal... ...expertise to support cutting-edge LLM research and production...Remote workContract workFor contractorsFreelance
$50 per hour
...This role focuses on improving and evaluating large language models through advanced mathematical reasoning... .... Contribute to new evaluation benchmarks spanning curricula from early undergraduate... ...to collaborate effectively in a remote setting. Self-motivated, able to...Remote workContract workFor contractorsFreelance- ...complex, multi-step ML tasks for frontier models. You will design experiments, implement... ...fall short. This full-time, fully remote role within the United States requires a... ...and strong Python/Git skills, with a focus on RL and LLM evaluation. #J-18808-Ljbffr ObsidianRemote jobFull timePart time
- ...researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating... ...written conclusions. You will work remotely in the United States for... ...teams through Cincinnatus’ extended workforce model. #J-18808-Ljbffr ObsidianRemote jobFull timePart time
- ...-based solutions, and perform rigorous data analysis. You will work remotely within the United States, about 35 hours per week, collaborating with researchers to ensure consistent, high-quality benchmark design and evaluation of frontier models. #J-18808-Ljbffr MercorRemote job
$70 per hour
...research problems that help evaluate advanced AI systems in scientific... ...tasks that strong frontier models do not consistently solve.... ...Contribute to a scientific computing benchmark by authoring rigorous... ...experience. Work Terms Remote, hourly engagement. Six-week...Remote workHourly payPart timeImmediate start- Scientist / Senior Scientist, Structure-Based Modeling Scientist / Senior Scientist, Structure-Based... ...Nice to have: Experience benchmarking across multiple PDB... ...for predictive modeling; Evaluate multiple structural representations... .... We offer: A remote-first team across the...Remote workFull timeFlexible hours
$30 - $50 per hour
...Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing AI research experiments,...Remote jobHourly pay$30 - $50 per hour
...projects. The role involves end-to-end research cycles, building and evaluating LLM systems, and collaborating on dataset development. The ideal... ..., and strong written communication skills. This is a remote, full-time position with competitive hourly compensation ranging...Remote jobHourly payFull time$70 - $85 per hour
...and geophysics challenges that evaluate whether advanced AI systems... ...focuses on creating original benchmark problems grounded in practical... ...problems with advanced AI models and refine them until they reach... ...terminal environment with remote compute sandboxes. Ability...Remote workHourly pay$80 - $135 per hour
...quality reference solutions for the CritPt benchmark (arXiv:2509.26574v3), a frontier... ...-verified reference data used to evaluate large language model performance on frontier physics reasoning... .... Work Terms Location: Remote. Employment type: hourly. Expected...Remote workHourly pay10 hours per week- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that frontier models cannot solve. Domains—depth in at least two subdomains with a...Part timeImmediate start
$50 - $70 per hour
...About the job Remote | Applied Physicist (AI Benchmarking) - $50-$70/hour We are sharing a specialised part-... ...in applied physics, mathematical modelling, experimental analysis, or computational... ...-choice assessment content, evaluate scientific accuracy and solution quality...Remote workFull timePart timeFor contractors10 hours per weekFlexible hours$100 - $120 per hour
...role combines end-to-end model development, from... ...training and TRADES. Evaluating robust accuracy under standard... ...parameter counts. LLM Post-Training and... ...evaluation, including benchmark construction, contamination... ...contributions. Work Terms Remote role. Hourly,...Remote workHourly payTemporary workFlexible hours- ...based AI development initiatives focused on enhancing frontier AI models. You will be responsible for identifying suitable mathematical... ...applied experience, and strong communication skills. This role is remote, offering an opportunity to work independently in a fast-paced...Remote work
- ...team combines deep expertise in model innovation and systems... ...We're looking for a Research Scientist who can define what "better"... ...'s model families, build the evaluation infrastructure to measure it... ...meaningful progress, not just benchmark performance. Develop novel quantitative...
- ...and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is perfect for someone with deep research expertise who...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote LLM Evaluation Scientist: Benchmarking Models. Be the first to apply!
- molecular biology scientist New York, NY
- water quality scientist New York, NY
- cosmetic scientist New York, NY
- machine learning scientist New York, NY
- principal applied scientist New York, NY
- image scientist New York, NY
- machine learning research scientist New York, NY
- hplc scientist New York, NY
- materials scientist New York, NY
- research associate scientist New York, NY



