ML Scientist, AI Evaluation & Benchmarking
Arena Intelligence, Inc.
A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'll engage in designing and analyzing comprehensive experiments while collaborating across teams. Applicants should possess a PhD in a relevant field and hands-on experience with large-scale models. The role offers competitive compensation and comprehensive benefits, fostering a culture centered around transparency and community impact. #J-18808-Ljbffr Arena Intelligence, Inc.
$196k - $230k
...AreNotion is the collaborative AI workspace where teams and... ...Researcher to define and scale how we evaluate Notion’s AI-powered... ...help teams spot regressions, benchmark improvements, and understand when... ...and working with Data Science/ML partners on measurement strategy...SuggestedLocal areaShift work$216k - $270k
Scale Labs, Research Scientist — AI Controls and MonitoringAs the leading data and evaluation partner for frontier AI companies, Scale... ...to establish standards and benchmarks for AI monitoring and escalation... ...experience addressing sophisticated ML problems, whether in a...SuggestedFull time$150k - $250k
...signal training data and evaluation infrastructure for frontier AI labs, with a founding team... ...that go beyond static benchmarks. Small team where individual... ...people · Industry: AI / ML — training data & evaluation... ...with the other Research Scientists to build shared...SuggestedFull timeVisa sponsorshipShift work$216k - $270k
Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale... ...to establish standards and benchmarks for AI monitoring and escalation... ...experience addressing sophisticated ML problems, whether in a...SuggestedFull time$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering with the world... ...to drive incremental improvements on benchmarks or optimize an existing process but instead... ...is measured. Researchers design evaluation frameworks that capture reasoning depth,...SuggestedWork at office3 days per week$180k - $260k
...About the RoleWe’re looking for an Applied Scientist, AI to turn messy, high-stakes healthcare... ...ll build strong baselines, design honest evaluations, run careful error analysis, and iterate... ...also be able to partner closely with ML engineering to productionize models, work...Temporary workWork at officeMonday to FridayMonday to Thursday$141.1k - $262.1k
...makes us Roche.Advances in AI, data, and computational... ...Intelligence (AI) to assist our scientists in both pRED and gRED to... ...-edge machine learning (ML) techniques. We are... ...objectives, training signals, and evaluation criteria.Evaluation & Benchmarks: Design and implement...Full timeWork experience placementLocal areaWorldwideRelocation package$167.4k - $310.8k
...makes us Roche.Advances in AI, data, and... ...Intelligence (AI) to assist our scientists in both pRED and gRED to... ...edge machine learning (ML) techniques. We are... ...training strategies, and evaluation methodologies.Model Capability... ...of rigorous reasoning benchmarks.You act as a technical...Full timeLocal areaWorldwideRelocation package$160k - $220k
...the RoleWe’re looking for an AI Research Scientist to advance the... ...architectures, new training or evaluation techniques, long-horizon research... ...validation standards are higher than benchmark culture alone, and you are... ...and engineers on rigorous ML research practices.External...Temporary workWork at officeMonday to FridayMonday to Thursday- ...applications, processes, and AI into a single, governed... ...AI Research Scientist to join our growing team... ...optimised RAG, tool‑use evaluation, and multi‑agent collaboration.Prototype and benchmark models; present findings... ...PhD in Computer Science, ML, or related field—or equivalent...Remote workFlexible hours
- DataAnnotation is seeking a Clinical Data Scientist for a remote contract role to evaluate AI-generated quantitative analyses and create benchmark problems for training AI systems. You will assess AI outputs, develop training problems across forecasting, experiment design...Remote jobContract work
$234.3k - $349k
...enterprises orchestrate AI-powered work. Our... ...world. As an AI research scientist, you'll be at the center... ...through model training, evaluation, and production deploymentDesign... ...novel evaluation benchmarks and methodologies that... ...7+ years of hands-on ML research experience, with...Full timeWork at officeLocal area- ...looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles... ...candidates have a PhD or equivalent experience in CS, ML, or related fields. Responsibilities include...Contract workSummer work
$84.13 - $91.34 per hour
AI Researcher - Efficient AI (Contractor) Step into the innovative... ...workflows. • Propose and evaluate novel compression methods (PTQ... ...vision, reasoning, and agentic benchmarks. • Contribute to publications,... ...or engineering experience in ML, efficient AI, model optimization...Full timeContract workTemporary workFor contractorsLocal areaImmediate start$200k - $325k
...anything. We're building the AI that finally changes that. Ivo... ...accurate on legal- specific tasks. Evaluate emerging work in agentic... ...Design and maintain datasets, benchmarks, and evals for training and measuring... ...at top venues — e.g., ML/AI conferences (NeurIPS, AAAI,...Contract workWork at officeImmediate startRemote workVisa sponsorshipRelocation packageFlexible hours- We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI... ...frameworks. Performance Evaluation Develop rigorous benchmarking methodologies for edge AI systems.... ...deployment frameworks such as: Core ML ExecuTorch ONNX Runtime LiteRT / TensorFlow...
$200k - $280k
...high‑performance computing for ML. Are comfortable working from... ...‑scale rollout collection and evaluation cheaper. Use these pipelines... ...as needed. Establish metrics, benchmarks, and experimentation... ...engineering. About Together AI Together AI is a research-driven...Full time- Carnaby Fox is seeking a Member of Technical Staff (AI Research) in San Francisco to help shape the... ...with world-class researchers to design experiments, evaluate LLMs, and improve data quality for high-stakes AI benchmarks. The role emphasizes independent ownership,...
- the company is seeking a Research Scientist to advance measurable recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and interpret... ..., and a track record in evaluating AI systems. #J-18808-Ljbffr United States...
- ...Department Technical About the Role Generative AI is transforming what's computationally... ...path through these bottlenecks. As an ML Research Scientist, you'll work at the frontier of... ...and likelihood estimation Develop and benchmark novel solver methods for diffusion ODEs...Full timeCasual workVisa sponsorship
- ...looking for an exceptional Research Scientist to develop next-generation AI technologies, focusing on user... ...applications. Design, prototype, evaluate, and deploy transformer-based generative... ...and establish reproducible benchmarking pipelines. Work closely with product...
- ...About the Role We’re looking for an Applied Scientist, AI to turn messy, high‑stakes healthcare... ...build strong baselines, design honest evaluations, run careful error analysis, and iterate... ...should also be able to partner closely with ML engineering to productionize models,...Temporary workWork at officeRelocation packageMonday to FridayMonday to ThursdayFlexible hours
- ..., we are building a team of world-class scientists, ML researchers, and engineers to work together... ...frontier of model architectures for AI x Chemistry: developing world models for... ...symmetries and constraints. Prototype, benchmark, and iterate rapidly to transform research...Work at office
- ...is hiring a Senior Machine Learning Scientist to lead autonomous life science systems... ...design architectures, workflows, and evaluation methods enabling AI to propose experiments, incorporate... ...-on role sits at the intersection of ML, biology, and automation, requiring translating...
- As a Research Scientist , you'll lead cutting-edge research that advances the state of generative AI for long-form storytelling. You'll work at the intersection... ...horizon generation. Build Novel Evaluation Frameworks Design robust benchmarks and evaluation methodologies for...Worldwide
- ...infrastructure layer for enterprise AI agents, and we are hiring a Research Scientist to advance the neuro-symbolic... ...Context Graph. Publish at top AI and ML conferences (NeurIPS, ICML, ICLR,... ...components with full observability and benchmarking. Engage with the Bay Area academic...Work at officeRelocation
- ...cutting-edge research into production-grade AI systems. You will design and build the... ..., including retrieval, orchestration, and evaluation, and you’ll take techniques from research... ...client-ready solutions. Ideal candidates ship ML-powered products end-to-end, balancing...
- Anthropic in San Francisco seeks a Research Scientist to measure recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and... ...RL, and policy teams to advance safe and reliable AI systems. #J-18808-Ljbffr Anthropic
$147k - $180k
Profluent is an AI‑first protein design company. Founded... ...Role We are expanding the ML Design Evaluation (MDE) program; the cross‑functional... .... We are looking for a scientist with deep expertise in... ...Contribute to campaign charters, benchmarking assay design, and post‑...- ...building quantum-accelerated AI servers to exponentially speed... ...We are looking for a Research Scientist who can help define Quantum AI... ...the intersection of frontier AI/ML, quantum algorithms,... ...test, and refine hypotheses. Benchmarking frameworks that reveal when a...Casual workVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Scientist, AI Evaluation & Benchmarking. Be the first to apply!
- scientist ii San Francisco, CA
- machine learning scientist San Francisco, CA
- scientist San Francisco, CA
- quality control scientist San Francisco, CA
- qc scientist San Francisco, CA
- regulatory scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- scientist antibody discovery San Francisco, CA
- applied scientist San Francisco, CA
- associate scientist San Francisco, CA
