Senior Research Scientist, Model Evaluation
Cohere
Who are we?Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must continue to develop new evaluation techniques that accurately reflect what models are already capable of, as well as set the agenda for what future models should be capable of. In this role, you are responsible for creating these next-generation evaluation methods and infrastructure to measure LLM progress.As a Senior Research Scientist, Model Evaluation, you will:Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish.Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations.Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency.Build scalable and reusable tools for digging into model performance.You may be a good fit if:You enjoy rapidly building prototypes that demonstrate the boundaries of what LLMs are capable of, and you have developed resources to measure those capabilities.You have spent dozens of hours reviewing complex data and LLM outputs to ensure high data quality.You are obsessive about rigorously measuring AI capabilities, and also about making sure your measurements actually align with the capabilities you care about.You have strong software engineering skills.Full-Time Employees at Cohere enjoy these Perks:A weekly lunch stipend of $75/75 or equivalent in your local currency for lunch.Full health and dental benefits, including a separate budget for mental health.RRSP matching, 401K, Pension Scheme.100% Parental Leave top-up for up to 6 months, for either parent.Annual enrichment benefits:Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.Education & learning stipend for conferences, courses, and coaching.6 weeks of paid vacation (30 working days!)Budget for traveling to other offices if you are remote, plus an annual company offsite.How and Where We Work:Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.For those not near an office: a co-working benefit so you can work alongside others in your city.Everyone receives a $500 home office stipend to set up your workspace properly.If any of the above doesn’t line up exactly with your experience, we still encourage you to apply. We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.LocationToronto; Canada; London; New York; San Francisco; Seattle; United StatesEmployment TypeFull timeLocation TypeHybridDepartmentModelingModeling
- Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New York | Negotiable | Unknown Why this role? Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must...SeniorFull timeWork at officeRemote workFlexible hours
$204k - $259k
Senior Research Scientist, Foundation Model (LLM/VLM) Waymo Position type: Full‑time Location: New York Waymo is an autonomous driving technology company... ...development Design and run experiments training and evaluating large deep learning models Present results to peers...SeniorFull time- Cohere is seeking a Senior Research Engineer, Model Evaluation, to create next‑generation evaluation methods and scalable infrastructure. You will develop benchmarks, datasets, and environments to measure frontier model capabilities, and you will push the state‑of‑the‑...Senior
- ...build cutting‑edge foundation AI models and end‑to‑end products that... .... Cohere is a team of researchers, engineers, designers, and more... ...Paris. Join us! Why this role? Evaluation is critical to making... ...measure LLM progress. As a Senior Research Engineer, Model Evaluation...SeniorFull timeWork at officeLocal areaRemote workHome office
- ...opportunities through their expert network. Qualified candidates should hold a MS or PhD in a relevant field and have experience in evaluating complex biology content. Strong communication skills and proficient English are essential for success in this role. #J-18808-...SuggestedImmediate start
- Cohere is seeking a Senior Research Scientist, Model Evaluation, to create ambitious evaluation benchmarks and scale evaluation infrastructure for enterprise AI. You will work with cross-functional teams to translate model feedback into trustworthy measurements and to push...SeniorRemote work
- Waymo is looking for a Senior Research Scientist for its Foundation Model team in New York. In this hybrid role, you will conduct applied research in deep learning to address autonomous driving challenges. The ideal candidate holds a Masters or Ph.D. and has over 3 years...Senior
- ...future team member for the role of SVP - Model Risk Management to join our Model Risk team... ...limitations clearly to stakeholders and senior management and partner stakeholders to... ...critical thinking skills, with the ability to evaluate complex model frameworks, identify risks,...SeniorWorldwide
- CDM Smith Inc. seeks a geologist/hydrogeologist to perform basic to moderate complexity evaluations of seismic hazards, slope stability, aquifer testing and hydrogeologic analyses for water supply and contaminant projects. Responsibilities include field data collection,...SeniorFor subcontractor
- ...future team member for the role of SVP - Model Risk Management to join our Model Risk team... ...limitations clearly to stakeholders and senior management and partner stakeholders to... ...critical thinking skills, with the ability to evaluate complex model frameworks, identify risks,...SeniorWorldwideFlexible hours
$124k - $280k
...Strategy& Strategy Consulting - Business Model Reinvention - Senior Manager you will provide strategic... ...in competitive analysis and market research- Applying systems thinking to identify... ...collaborating closely with team members. We evaluate these factors thoughtfully to...SeniorFull timeH1b$77k - $202k
...Strategy& - Strategy Consulting Business Model Reinvention - Senior Associate, you will provide strategic... ...competitive analysis and market research to inform strategic planning- Utilizing... ...closely with team members. We evaluate these factors thoughtfully to establish...SeniorFull timeH1b$228.7k - $343.1k
...financial crime at enormous scale, and one bad model can mean millions in credit losses,... ...team validate at scale, so you critically evaluate what it produces and own the evaluation... ...figured out how to oversee yet. As a senior individual contributor, you lead through...SeniorRemote jobFull timeLocal areaShift work- ...providing independent assurance and evaluating the company's risk management... ....We are looking for data scientists and AI developers who will... ...ResponsibilitiesProficiency in frameworks for auditing models, including criteria like... ...Engineering, Operations Research, or Economics.- Minimum of 5...Senior
$203k - $338.3k
Position Summary Regulatory & Financial Risk - Senior Manager - Model Validation Our Deloitte Regulatory, Risk & Forensic team helps client... ...and stress testing, error analysis, and scenario-based evaluation).Develop models (example: credit risk models) and...Senior$184.9k - $217.5k
...make an impact?West Monroe is seeking a Senior Manager to join our growing Office of the... ...automation into an integrated operating model for the Office of the CFO.As a Senior Manager... ...definition, location strategy, provider evaluation, transition planning, cutover, hypercare,...SeniorWork at officeLocal areaImmediate startFlexible hours$192k - $304.75k
...computing. Today, we are building software, systems, and research platforms that help scientists and engineers solve problems that were once out of... ...precision algorithms, and sparse direct/iterative hybrids.Evaluate algorithms on workloads in mechanics, contact, thermal-...SeniorFull timeRemote work- ...for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client lab as part of the extended...Weekday work
- YO IT Consulting is seeking a Senior Accountant to evaluate AI-generated financial calculations and ensure accuracy in accounting practices. With... ...accounting experience, you will work remotely and challenge AI models on various real-world accounting scenarios. The ideal...SeniorRemote job
$166k - $200k
...We are seeking Research Scientists to lead deep, foundational AI research and post-research implementation that powers our... ...), reinforcement learning systems, and comprehensive evaluation frameworks across models, inputs, and outputs. This is not just a feature-iteration...SeniorTemporary workWork at officeRemote workShift work$200k - $320k
...strengthen our expansive research team. We are looking... ..., and computer scientists—who are passionate about... ...learning and large language models, develop predictive... ...data, and more. As a Senior Research Scientist you... ...prototypes, run large‑scale evaluations, and drive the project...SeniorWork at officeLocal areaRelocation$113.87k - $165.11k
...the Opportunity The Kostas Research Institute (KRI) at... ...relationship with and support of KRI Senior R&D Engineers/Scientists for government and... ...integrating simulation‑based models, including physics‑based,... ...model training, fine-tuning, evaluation, and experimentation....SeniorWork experience placementWork at office$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs...Full time- Vibrant Emotional Health is seeking a Lead Research Scientist to provide scientific leadership for a high-impact research and evaluation portfolio supporting the 988 Lifeline and related activities. You will guide analytic planning, oversee literature reviews, and serve...SeniorRemote job
- We are seeking an expert to evaluate and improve our AI models through comprehensive testing and analysis. You will be responsible for designing evaluation frameworks, conducting model assessments, and providing actionable insights for model improvement. Key Responsibilities...
- Mercor is seeking a Generalist who can operate in English and Punjabi. This contract, remote position focuses on evaluating AI outputs and supporting model evaluation tasks. You will conduct fact-checking, assess reasoning, clarity, tone and completeness, and provide actionable...Remote jobContract work
- Cincinnatus LLC is hiring for a Marketing SME to support GenAI model evaluation and brand/growth tasks. The role centers on applying rigorous marketing judgment to AI training data and guiding cross-functional teams to improve model outputs. Ideal candidates have extensive...
$20 per hour
...creative and technical talent with leading AI research labs. Headquartered in San Francisco,... ...tools. Generate high-quality human evaluation data by identifying response strengths,... ..., and completeness of responses. Ensure model responses align with expected conversational...Remote jobContract workPart timeSummer work- Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts...Remote job
- AuraOne is seeking an Evaluation Harness Model Evaluation Specialist to work remotely as an independent contractor. You will review evaluation harness model outputs, label issues, and provide structured feedback to retrain the model using AuraOne's quality rubric. Strong...Remote jobFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Research Scientist, Model Evaluation. Be the first to apply!
- molecular biology scientist New York, NY
- water quality scientist New York, NY
- cosmetic scientist New York, NY
- machine learning scientist New York, NY
- principal applied scientist New York, NY
- image scientist New York, NY
- machine learning research scientist New York, NY
- hplc scientist New York, NY
- materials scientist New York, NY
- research associate scientist New York, NY



