ML Scientist, AI Evaluation & Benchmarking
Arena Intelligence, Inc.
A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'll engage in designing and analyzing comprehensive experiments while collaborating across teams. Applicants should possess a PhD in a relevant field and hands-on experience with large-scale models. The role offers competitive compensation and comprehensive benefits, fostering a culture centered around transparency and community impact. #J-18808-Ljbffr Arena Intelligence, Inc.
- ...Fleet AI, Inc. is seeking a Research Scientist to join their core research team in San Francisco. This role focuses on investigating how environments... ...labs. Key responsibilities include generating benchmarks to evaluate frontier models, automating environment...Suggested
$167.4k - $310.8k
...us Roche. Advances in AI, data, and computational... ...(AI) to assist our scientists in both pRED and gRED to... ...edge machine learning (ML) techniques. We are seeking... ...strategies, and evaluation methodologies. Model Capability... ...of rigorous reasoning benchmarks. You act as a...SuggestedLocal areaWorldwideRelocation package$196k - $230k
...Notion is the collaborative AI workspace where teams and agents... ...Researcher to define and scale how we evaluate Notion’s AI-powered... ...help teams spot regressions, benchmark improvements, and understand when... ...and working with Data Science/ML partners on measurement strategy...SuggestedLocal areaShift work$216k - $270k
Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale... ...to establish standards and benchmarks for AI monitoring and escalation... ...experience addressing sophisticated ML problems, whether in a...SuggestedFull time- ...looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles... ...candidates have a PhD or equivalent experience in CS, ML, or related fields. Responsibilities include...SuggestedContract workSummer work
- About Thorin Thorin is an applied AI company born out of 8VC’s... ...architectures, training paradigms, and evaluation techniques tailored to... ...Design, implement, and test ML / AI methods that improve model... ...appropriate (writing, talks, shared benchmarks). What You’ll Bring Advanced...
- ..., we are building a team of world-class scientists, ML researchers, and engineers to work together... ...frontier of model architectures for AI x Chemistry: developing world models for... ...symmetries and constraints. Prototype, benchmark, and iterate rapidly to transform research...Work at office
- ...building quantum-accelerated AI servers to exponentially speed... ...We are looking for a Research Scientist who can help define Quantum AI... ...the intersection of frontier AI/ML, quantum algorithms,... ...test, and refine hypotheses. Benchmarking frameworks that reveal when a...Casual workVisa sponsorship
$147k - $180k
Scientist II, ML - Guided Protein Design Evaluation Emeryville, California, United States; Hybrid (2-3 days on‑site) Profluent is an AI‑first protein design company. Founded in 2022, we develop deep... ...to campaign charters, benchmarking assay design, and post‑campaign...- Knowtex is seeking an ML Scientist (Research) to enhance its voice AI and clinical NLP capabilities within healthcare. You will develop and evaluate innovative machine learning solutions aimed at advancing medical speech recognition and clinical language understanding....
- ...Department Technical About the Role Generative AI is transforming what's computationally... ...path through these bottlenecks. As an ML Research Scientist, you'll work at the frontier of... ...and likelihood estimation Develop and benchmark novel solver methods for diffusion ODEs...Full timeCasual workVisa sponsorship
$141.1k - $262.1k
...makes us Roche. Advances in AI, data, and computational... ...(AI) to assist our scientists in both pRED and gRED to... ...cutting‑edge machine learning (ML) techniques. We are... ...objectives, training signals, and evaluation criteria. Evaluation & Benchmarks: Design and implement...Work experience placementLocal areaWorldwideRelocation package- ...mining to translate scientific insights into actionable AI solutions. Conduct experiments, evaluate model performance, and iterate rapidly to improve... .... Hands‑on experience building and deploying scalable ML models on cloud platforms (e.g., AWS). Demonstrated ability...
- ...infrastructure layer for enterprise AI agents, and we are hiring a Research Scientist to advance the neuro-symbolic... ...Context Graph. Publish at top AI and ML conferences (NeurIPS, ICML, ICLR,... ...components with full observability and benchmarking. Engage with the Bay Area academic...Work at officeRelocation
$160k - $280k
...time. We are a team of musicians and AI experts, including alumni from... ...build and deploy our state of the art ML models trained with an H100/scientist ratio of >100x. Check out our Suno... ...engineering, designing, training and evaluating machine learning models Track...Work at officeFlexible hours- ...About the role AI research at WRITER isn't just... ...world. As an AI research scientist, you'll be at the center... ...model training, evaluation, and production deployment... ...Build novel evaluation benchmarks and methodologies that... ...need 7+ years of hands‑on ML research experience, with...Full timeLocal areaFlexible hours
- At Goaly, our mission is to make custom AI affordable for every business. Our... ...at scale. Design domain‑tailored eval benchmarks: Build evaluation frameworks that capture real‑world performance... ...on Hugging Face or competitive ML achievements (Kaggle medals, competition...Full timeWork at office
$70 - $100 per hour
Mercor is seeking a STEM Computational Scientific Software & Evaluation Design specialist to tackle complex computational problems remotely... ...scientific software libraries. You will design challenges to assess AI models and collaborate with research teams to refine these...Remote jobFlexible hours- ...Role We’re looking for an AI Research Scientist to advance the methodological... ...architectures, new training or evaluation techniques, long-horizon... ...validation standards are higher than benchmark culture alone, and you are... ...and engineers on rigorous ML research practices....Full timeTemporary workWork at officeMonday to FridayMonday to ThursdayFlexible hours
- ...Intelligence is the open platform for evaluating how AI models perform in the real... ...variety of Machine Learning Scientist to help advance how we... .... Strong foundation in ML and statistics, with a track... ...that go beyond traditional benchmarks Analyze large-scale human voting...Permanent employmentWork at office
- ...are partnering with a AI‑native therapeutics company... ...a Machine Learning Scientist to conduct original, high... ...implement experiments, evaluate results, and clearly... ...modalities Define meaningful benchmark tasks and evaluation... ...scientists, and ML researchers Evaluate modern...
- ...Blank Bio is an applied AI research lab focused on... ...a technical team of AI scientists and engineers from... ...Institute. The Role As an ML research scientist, you... ...ML methods and develop benchmarks that reflect clinically... ...Develop benchmarks and evaluation frameworks for tasks spanning...
- ...further notice. Meet the Team At Foundation AI, we are leading frontier AI research... ...training methods, inference optimization, evaluation techniques, and data pipelines. Together,... ...software engineering practices, and common AI/ML libraries Experience with large language models...
- ...systems / high‑performance computing for ML. Are comfortable working from... ...make large‑scale rollout collection and evaluation cheaper. Use these pipelines to train, evaluate... ...and APIs as needed. Establish metrics, benchmarks, and experimentation frameworks to validate...Full time
- ...creating trustworthy and reliable AI systems, changing banking for... ..., our applications of AI & ML bring humanity and simplicity... ...cross‑functional team of data scientists, software engineers, machine learning... ...from design through training, evaluation, validation, and...Flexible hours
- Patronus AI is seeking an Applied Researcher in San Francisco to lead foundational research... ...how agentic AI systems are trained, evaluated and improved. You will work at the intersection... ...research. A BS, MS, or PhD in CS/ML is required; strong Python and ML framework...
- ...consumer-grade agents that redefine human-AI collaboration for millions. Software shouldn... ...our research rigor, safety metrics, and evaluation pipelines for everything that calls itself... ...Qualifications PhD or equivalent track record in ML / NLP / RL / systems 3+ strong papers or...Relocation package
- ...software that operationalizes responsible AI governance at scale. We're a 4-month-old AI... ...AI and agentic systems Develop risk evaluation methodologies that adapt as threats evolve... ...cybersecurity, with 2+ years focused on AI/ML security, red teaming, or adversarial testing...Part timeRemote workFlexible hours
$50 per hour
A leading AI research firm is seeking PhDs in Mathematics or related fields for a fully remote contract role. The successful candidate... ...will design advanced math problems to test AI performance and evaluate outputs for accuracy. Strong mathematical reasoning, problem-...Remote jobContract workFlexible hours- AI Researcher (Computer Vision/Multimodal/Generative AI) About the Role We are hiring ML Researchers to develop novel approaches that advance the frontier of multimodal vision... ...aligned with product differentiation. Evaluate new model paradigms for scalability and efficiency...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Scientist, AI Evaluation & Benchmarking. Be the first to apply!
- research associate scientist San Francisco, CA
- scientist antibody discovery San Francisco, CA
- protein scientist San Francisco, CA
- lead scientist San Francisco, CA
- applied sports scientist San Francisco, CA
- deep learning scientist San Francisco, CA
- lab scientist San Francisco, CA
- principal scientist San Francisco, CA
- molecular biology scientist San Francisco, CA
- senior analytical scientist San Francisco, CA

