Frontier AI Evaluation Scientist
Neura Market
OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model environments and steer major training runs so that research outputs translate into real-world products. Collaborate with researchers, engineers, product and safety teams to decide what to measure, how to measure it, and how to ship improvements. This high-autonomy role rewards curiosity and clear communication. #J-18808-Ljbffr Neura Market
$216k - $270k
Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you will design testing harnesses, collaborate with agencies, and publish reports to inform policymakers...Suggested- An innovative tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and...Suggested
- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models through structured technical assessments and focus on realistic data engineering workflows. The role involves reviewing model-...Suggested
- A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models and AGI/ASI measurement. Candidates should have strong... ...AI measurements and build evaluation environments that drive progress. #J...Suggested
- ...is seeking exceptional research engineers to push the boundaries of frontier AI safety, shaping empirical understanding of risk and owning end-to-end threads within this effort. You’ll design evaluations of frontier models against real threat models, develop datasets,...Suggested
$350k
Mirendil is a tech-first company located in San Francisco, California, seeking a Research Engineer to develop evaluation infrastructure for AI models. You'll design frameworks that measure model capabilities, implement automated pipelines, and create workflows for inspecting...- Ent is seeking a Frontier Research AI Scientist to lead cutting-edge AI/ML research and develop novel workflows for security problems. The role invites deep technical breadth and the ability to combine techniques to create new frameworks, while contributing to a security...Remote jobWork at office
- Sygaldry is building quantum-accelerated AI systems and seeks a Research Scientist to define Quantum AI beyond quantum machine learning, exploring how... ...inference, and control. You will work at the intersection of frontier AI/ML, quantum algorithms, and hardware-software co-...
- Snorkel AI in San Francisco is searching for a Research Scientist to lead the development of datasets and benchmarks for AI models. This customer-facing role... .... in a relevant field and a strong focus on AI/ML evaluation and dataset design. With robust growth opportunities...
$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering with the world’... ...rearchitect critical operations for the frontier of AI. Our customers include the largest... ...ForAt Distyl, we build AI systems using Evaluation-Driven Development—an approach where evaluation...Work at office3 days per week- ...conduct impactful research to improve the security and privacy of frontier intelligence systems. Responsibilities include developing threat... ...role offers opportunities to drive meaningful security improvements in AI systems. #J-18808-Ljbffr United States Digital Space LLC
- ...DesignArena, invites you to join a talent-dense team in San Francisco with 5.5M+ users and a rapidly growing platform. You will define how frontier AI models are measured, design new benchmarks, run experiments, and publish analyses that become industry gold standards. We sponsor...RelocationVisa sponsorship
$211k - $290.5k
...Compass — Faire’s user facing AI bet within the Discovery... ...behalf.As a Senior Applied AI/ML Scientist on the Compass team, you will... ...agent quality through data, evaluation, and modeling, while shipping... ...product fast.You will work at the frontier of agentic AI, blending...Work experience placementWork at officeLocal areaImmediate startRemote workMonday to FridayFlexible hours3 days per week- Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems...
$160k - $220k
...multi-year runway.About the RoleWe’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare. This role is ideal for... ...may include novel architectures, new training or evaluation techniques, long-horizon research bets, peer-reviewed...Temporary workWork at officeMonday to FridayMonday to Thursday$234.3k - $349k
...leading enterprises orchestrate AI-powered work. Our vision is... ...the world. As an AI research scientist, you'll be at the center of... ...hypothesis through model training, evaluation, and production... ...— representing WRITER at the frontier of the field and contributing...Full timeWork at officeLocal area- ...candidate to conduct critical comparative analysis to advance our understanding of model capabilities. You will build and refine evaluation systems that create tight feedback loops between data, evals, and model behavior, and develop generalizable frameworks for reasoning...
$130k - $220k
...Artificial Analysis is the leading independent AI benchmarking and insights company. They... .... Their benchmarks do not just measure frontier AI — they actively shape it. The company... ...** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not...Full timeWorldwide- ...organize human intelligence to power the AI economy. We partner with leading AI labs... ...development. Our vast talent network trains frontier AI models in the same way teachers teach... ...As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of...Full timeWork at officeRelocation package
- Heyaristotle in San Francisco is looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles into actionable measures. The role starts as a summer contract with potential...Contract workSummer work
- ...Mistral Mistral provides full‑stack AI solutions: from frontier models to developer tools, applications... ...team‑spirited. The Role As an AI Scientist , you will research and develop novel... ...tooling and infrastructure to train, evaluate, and analyse AI models at scale. Your...Relocation package
$1,750 - $2,150 per month
Obsidian is looking for experienced cybersecurity professionals to review AI systems' threat detection and vulnerability assessments. Responsibilities include evaluating AI outputs and creating realistic cybersecurity scenarios. Ideal candidates should have over 3 years...- About: Frontier AI x Biology | Foundation Models | Therapeutic Discovery Stage: Well-funded... ...biologists and experimental scientists to develop the next generation of biological... ...model performance through post-training, evaluation, alignment and fine-tuning techniques Work...
$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators who compare text and voice samples to assess naturalness and authenticity. You will listen to audio clips, rate quality, and flag any mismatches in tone or pronunciation, with emphasis on cultural context...Remote jobFlexible hours$400 per month
Obsidian is seeking contributors for a project with a leading AI research lab focused on evaluating frontier AI coding models. You'll complete and assess complex data engineering tasks using AI coding agents. The role requires at least 2 years of experience in data engineering...Remote job- ...AI Research Scientist (Robot Learning) San Francisco AI & Software In office Full-time... ...Engineer (Robot Learning) you will drive frontier AI model development and data flywheel... ...collection, model training and model evaluation in the real world. As an early...Full timeWork at officeImmediate start
- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes, and...
$400 per month
About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering...- Sprinter Health seeks an AI Research Scientist to push the methodological frontier of AI in healthcare. You will develop and own a research agenda aligned with... ..., exploring novel architectures, training and evaluation methods, and validation studies that can graduate...Flexible hours
- Welo Data is looking for a Data Labeling Associate in San Francisco to evaluate AI systems' handling of Arabic nuances. The role requires professional-level proficiency in Arabic and 2 years of AI safety experience. Responsibilities include critiquing Arabic AI outputs...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Frontier AI Evaluation Scientist. Be the first to apply!


