Machine Learning Research Scientist, Evaluations
$180.6k - $225.75kScale AI
Scale works with the industry's leading AI labs to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research Organization and will focus on building benchmarks and diagnosing model failure modes in both text and multimodal modalities.In this role, you will develop rigorous evaluations and diagnostic methods that reveal where frontier models fail and why. You will collaborate with researchers and engineers to define best practices in evaluation-driven AI development. You will also partner with top foundation model labs to translate failure analysis into technical and strategic input on the next generation of generative AI models.You will:Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA.Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities.Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them.Publish research findings in top-tier AI conferences.Ideally you’d have:Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.Excellent written and verbal communication skills.Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals.Previous experience in a customer facing role.Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend.Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is:$180,600—$225,750 USDPLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants.About Us:At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst& Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. We are expanding our team to accelerate the development of AI applications.We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status. We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at View email address on click.appcast.io. Please see the United States Department of Labor's Know Your Rights poster for additional information.We comply with the United States Department of Labor's Pay Transparency provision. PLEASE NOTE: We collect, retain and use personal data for our professional business purposes, including notifying you of job opportunities that may be of interest and sharing with our affiliates. We limit the personal data we collect to that which we believe is appropriate and necessary to manage applicants’ needs, provide our services, and comply with applicable laws. Any information we collect in connection with your application will be treated in accordance with our internal policies and programs designed to protect personal data. Please see our privacy policy for additional information.
$30 - $50 per hour
A tech company is seeking an AI Researcher to support end-to-end research for modern AI... ...involves designing experiments, defining evaluation protocols, and improving evaluation... ...a fundamental understanding of AI and machine learning, along with experience in NLP or computer...SuggestedRemote jobHourly pay- Aaru is seeking an Evaluation Researcher to tackle challenging measurement problems at the intersection of machine learning and behavioral science. In this role, you will design studies, build tests, and communicate evidence to impact decision-making. Located in New York...Suggested
$180.6k - $225.75k
...and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with... ...Master's degree in Computer Science, Machine Learning, AI, or a related field.Deep... ...us to ensure a fair and thorough evaluation of all applicants.About Us:At Scale...SuggestedFull time$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding... ....A track record of published research in machine learning, particularly in generative AI.At least...SuggestedFull time$150k - $200k
ML Research Scientist -Deep Learning & Transformer ArchitecturesPlease direct all resume submissions to QuantTalentUS... ...candidate will have a PhD in machine learning or a related field and... ...multi-scale representations• Build evaluation frameworks for next-token prediction...Suggested$85k - $150k
...Research Scientist, Artificial Intelligence (PhD) New York City About Synaptrix... ...the right fusion of deep learning, signal processing, and... ...deep technical expertise in Machine Learning, Artificial... ...reproducible experiments and evaluation. Strong mathematical foundations...Full time- ...A health tech company in New York City is hiring a Research Scientist to evaluate the impact of ambient AI on healthcare outcomes. The role emphasizes designing studies, engaging with health systems, and fostering collaboration across product teams. A PhD in a relevant...Work at office
$320k - $400k
As a Research Scientist on our team, you will partner with Research Engineers... ...foundation models that learn the joint dynamics of distributed... ..., RL training loops, and evaluation infrastructure needed to train... ...in generative AI and machine learning, building specialized...- ...customers. Cohere is a team of researchers, engineers, designers, and... .... Join us!Why this role?Evaluation is critical to making progress... ...progress.As a Senior Research Scientist, Model Evaluation, you will:... ...credit.Education & learning stipend for conferences, courses...Full timeWork at officeLocal areaRemote workHome office
- ...what it means to reason, to learn, to make decisions, to understand... ...first. About the Role Research scientists lead Basis’ efforts to develop... ..., programming languages, machine learning, computational neuroscience... ...materials for both hiring evaluation and recruitment-related...Full time
$174k - $252k
Senior Research Scientist, Generative Media, Apparel ML Share Senior Research Scientist, Generative... ...with C++, Python, generative AI, and machine learning. One or more scientific publication... ...experiences. Implement, train, evaluate, and iterate on machine learning models...Temporary workWorldwide- ...re driven by continuous learning, rapid pivots, and the... ...Overview We are expanding our research team and looking for an experienced Research Scientist specializing in Conversational AI & Machine Learning. In this role,... ..., and experimental evaluation. Ideally with a focus...Remote work
- OpenRouter in New York seeks a Research Scientist to advance how the world understands, evaluates, and routes large language models. You will design experiments, build evaluation frameworks, and publish findings that influence rankings and routing decisions. You will collaborate...
- Scale Labs, Research Scientist — Frontier Risk Evaluations As the leading data and evaluation partner for frontier AI companies, Scale plays an integral... .... A track record of published research in machine learning, particularly in generative AI. At least three years...
- ...growing team of practicing MDs, AI scientists, PhDs, creatives,... ...The Role Abridge is hiring Research Scientists to join our Strategic... ...Research team to rigorously evaluate and advance the real-world impact... ...proactive mindset, with a desire to learn and grow as a researcher in a...Hourly payFull timeWork at officeRelocation packageFlexible hours
- ...Job Title AI Research Scientist Location Hybrid / Remote Employment Type Full... ...research in artificial intelligence, machine learning, and generative AI. The ideal... ...generative AI. Design, develop, and evaluate client AI models, algorithms, and architectures...Full timeRemote work
- ...team.KPMG is currently seeking an AI Research Scientist to join our Audit Technology Alliance... ...research agenda by implementing and evaluating new techniques from academic literature... ...science role with a focus on NLP, machine learning, or a related fieldBachelor's degree...Work experience placementLocal area
$90 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ..., Larry Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation...Contract workSummer workRemote work- As a Research Scientist on our team, you will partner with Research Engineers... ...foundation models that learn the joint dynamics of distributed... ..., RL training loops, and evaluation infrastructure needed to train... ...in generative AI and machine learning, building specialized...
- Fleet AI, Inc. seeks an Applied Research Scientist to collaborate with technical teams and deliver... ...role demands strong engineering and machine learning skills to analyze requests and ensure the delivery of artifacts ready for evaluation. Ideal candidates will have...
- Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New York | Negotiable | Unknown Why this role? Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must...Full timeWork at officeRemote workFlexible hours
- Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross...Remote jobHourly payFlexible hours
$30 - $50 per hour
A tech company specializing in AI research is seeking a mid-senior level researcher to manage applied AI research projects. The role involves end-to-end research cycles, building and evaluating LLM systems, and collaborating on dataset development. The ideal candidate should...Remote jobHourly payFull time$120k - $170k
...software. WhatYou’llBeLaunching As an AI Research Scientist, you will join the AI and Innovation... ...&D)ofAI systems, models,and advanced machine learning algorithmsthataugment physics-driven... ...(connectors, simulation frameworks, evaluation harnesses) to refine simulation-based...Full timeWork experience placementCurrently hiringRemote work$216k - $270k
Research Scientist, AI Controls and Monitoring Scale Labs, Research Scientist - AI Controls and... ...Monitoring As the leading data and evaluation partner for frontier AI companies, Scale... ...record of published research in machine learning, particularly in generative AI. At least...Full time$204k - $259k
Senior Research Scientist, Foundation Model (LLM/VLM) Waymo Position type: Full‑time Location... .... The Applied Research team develops machine learning solutions for autonomous driving... ...Design and run experiments training and evaluating large deep learning models Present results...Full time$218.4k - $273k
...AGI), and building upon our prior model evaluation work with enterprise customers and... ...Environments (ACE) team, part of Scale’s Research organization, brings together customer... ...stack (eg. AWS or GCP) and developing machine learning models in a cloud environment.Our...Full time$200k - $300k
...infrastructure to provide liquidity to the global markets.As a Machine Learning Researcher at Virtu, you'll pursue high-impact research opportunities... ...of what industry you come from. The RoleInvestigate, evaluate, and prototype innovative algorithmic solutions using novel...$200k - $350k
...committed to world-class research. We empower... ...leading reinforcement learning research and trading at... ...exceptional Research Scientist/Research Engineer to join... ...trading: designing and evaluating policy architectures,... ...in Computer Science, Machine Learning, Robotics (or...$141.1k - $262.1k
...discovery and development. Roche’s Research and Early Development... ...(AI) to assist our scientists in both pRED and gRED to deliver... ...discovery with cutting-edge machine learning (ML) techniques. We are seeking... ..., training signals, and evaluation criteria.Evaluation &...Full timeWork experience placementLocal areaWorldwideRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Research Scientist, Evaluations. Be the first to apply!
- molecular biology scientist New York, NY
- water quality scientist New York, NY
- cosmetic scientist New York, NY
- machine learning scientist New York, NY
- principal applied scientist New York, NY
- image scientist New York, NY
- machine learning research scientist New York, NY
- hplc scientist New York, NY
- materials scientist New York, NY
- research associate scientist New York, NY


