Frontier Language Model Evaluation Engineer
Artificial Analysis, Inc.
Artificial Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets, scoring systems, and evaluation infrastructure applicable across major models released by labs worldwide. Based in San Francisco or flexible to other major cities, this role offers equity and the opportunity to shape how industry measures frontier AI capability while collaborating with leading researchers and #J-18808-Ljbffr Artificial Analysis, Inc.
- ...benchmarking company. We support labs, engineers and enterprises to understand AI... ...of AI, they are actively shaping the frontier. Our benchmarks and analysis are trusted... ...other industry leaders. The Opportunity Language model evaluation is the sharpest question in AI: what...Language
$180k - $225k
...transformation through frontier AI systems that solve... ...ever. New foundation models, reasoning techniques,... ...remains one of the hardest engineering challenges.As a... ...customers to design, evaluate, and deploy intelligent... ...latest advances in large language models, reasoning, retrieval...LanguageFull time$98k - $140k
...products. You’ll work with product and engineering teams to build systems to define... ...context engineering, designing evaluation systems, and analyzing data. This... ...of that you'll shape Notion's model strategy and work directly with frontier AI labs (OpenAI, Anthropic, Google...SuggestedLive inLocal area- ...Technical Staff to build the next generation of AI benchmarks. You will design frontier evaluations, construct evaluation datasets, and publish analyses that set industry standards for model capability and reliability. You will work across leading labs and enterprises,...Suggested
- ...remotely. The role focuses on fact-checking and generating evaluation data, requiring native fluency in Urdu and strong English writing... ...a bachelor’s degree and significant experience with large language models. Responsibilities also include independently assessing...LanguageRemote job
$238k - $302k
...simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language Models (LLMs) and Vision-Language... ...are looking for quantitatively-minded engineers to research and propose new ways to...LanguageFull timeRemote work$15 - $20 per hour
...tools . Generate high-quality human evaluation data by identifying response strengths,... ...and completeness of responses. Ensure model responses align with expected conversational... .... Significant experience using large language models (LLMs). Excellent writing...LanguageContract workSummer workRemote work- Mercor is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab.... ...of music in your bilingual language. Ideal candidates have 3+ years in music production or audio engineering, a degree in music, and fluency in French...LanguagePart timeImmediate start10 hours per week
- A cutting-edge AI firm in San Francisco seeks a VLM Post-Training Owner to lead enterprise engagements and enhance vision-language models. The role combines project ownership with technical execution, ensuring quality data generation and customer satisfaction in AI solutions...Language
- Mercor is seeking experienced musicians to evaluate generative musical AI models in collaboration with a leading AI lab. You will assess model outputs across categories of music in your bilingual language, focusing on lyrics, voice generation, and audio quality. The role...LanguagePart time
- Welo Data in San Francisco seeks a full-time AI Evaluator with professional proficiency in Portuguese (Portugal) and experience in Generative AI safety. The role involves critiquing AI outputs, identifying biases, and refining evaluation frameworks. Candidates should possess...LanguageFull time
- ...is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab... ...other standards, using your bilingual language skills. Ideal candidates have 3+ years as a music producer or audio engineer, a college degree in music, and...LanguagePart timeImmediate start10 hours per week
$220k - $320k
...net]( trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the... ...-to-end: distillation, training, evaluation, and planet-scale hosting. We... ...a well-funded ten-person team of engineers who work in-person in downtown San...LanguageFull timeWork at office- ...is seeking experienced musicians to evaluate generative musical AI models in collaboration with a leading AI... ...categories of music in your bilingual language and contribute to structured... ...years in music production or audio engineering, a music degree, and native or near...LanguagePart timeImmediate start10 hours per week
- ...standards, and research advanced defenses against emerging threats. Applicants should have deep expertise in LLM safety, strong software engineering skills, and relevant academic qualifications in AI or related fields. This position is pivotal for paving the way for safe AI...
$110k - $130k
...AI teams train and run models on the right data. Our... ..., annotates, and evaluates data across the full AI... ...of 100+ working at the frontier of AI and have raised... ...As a Customer Support Engineer, you will be working with... ...with SQL and scripting languages such as Python or...Language- ...RoleAs a Forward Deployed Engineer you will work closely... ...wants to be at the frontier of enterprise AI adoption... ...solutions.Evaluate build-vs-buy options for... ...using modern programming languages such as Python, JavaScript... ...platforms, enterprise data models, or regulated...LanguageFull timeTemporary workLocal areaImmediate start
$155.4k - $233.2k
...services operate. By combining frontier agentic AI, an enterprise-... ...bar behind Harvey's human evaluations. As we scale globally, the volume... ...enough that Product, Engineering, and AI Research act on it to... ...quality across multiple markets, languages, or jurisdictions Built...LanguageContract work$80 per hour
...and Jack Dorsey . Position: Data Engineer (Coding Agent Experience) Type: Contract... ...Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations involving...Contract workSummer workRemote work$218.5k - $288k
...Applied Scientist specializing in Small Language Models and AI Training, you will lead... ...You will work closely with research, engineering, and product teams to advance model training... ...language models.Design, implement, and evaluate model training experiments to improve...LanguageWork at officeFlexible hours3 days per week$315k
We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand... ...Safety Level (ASL) to assign to models. Research leads on this team... ...workstreams, we would value experience with language model agents, although this is not...LanguageCurrently hiringWork at officeImmediate startHome officeVisa sponsorshipRelocation package- Engineering Manager, Foundation Model Inference (FMAPI)RDQ427R519At Databricks, we are driven by a passion to... ...to serve, scale, and optimize frontier models with enterprise-grade reliability... ..., gender identity or expression, language, national origin, physical and mental...LanguageWorldwide
$120k - $200k
...The Role As a Founding Engineer, you will work... ...and precedent Create evaluation infrastructure for high... ...system architecture and model behavior to UI, deployment... ...deeply curious about frontier multimodal and agentic... ...infrastructure, and modern language and vision models. How...LanguageFor contractorsLocal area- Obsidian is seeking a Spanish Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project. You will handle... ..., and evaluation tasks to help train and benchmark advanced language models. The ideal candidate should have strong writing skills,...LanguagePart time10 hours per week
- ...software and foundation models enable vehicles to... ...autonomy. Working at the frontier of a still-undefined... ...collaborate with research and engineering teams to define what “... ...pretraining and evaluation Manage and mentor a... ...datasets involving video, language, lidar, radar and...LanguageFull timeWork at officeRemote workWork from homeVisa sponsorshipRelocation packageFlexible hours
$300k - $320k
...Program Manager to lead our AI model evaluation initiatives across multiple... ...Research, Trust & Safety, Frontier Redteaming, and Policy... ...programs in AI development, ML engineering, or related fields. You’ll... ...prompt engineering on language models Have experience designing...LanguageWork at officeHome officeVisa sponsorshipRelocation package$134.5k - $265.1k
...At Deloitte, Forward Deployed Engineers (FDE) don’t just build AI... ...0/30/2026.Work you’ll doAs a Frontier GenAI FDE, you will work side... ..., safety, latency, cost, and model risk.Deliver production-quality... ...with MLOps/LLMOps practices: evaluation frameworks, model monitoring,...Local area- Welo Data is seeking Data Labeling Associates in San Francisco to evaluate AI systems focused on Arabic language nuances. The role includes model evaluation, identifying biases in datasets, and ensuring quality across global AI workflows. Candidates should have native French...
$155.6k - $306.8k
...At Deloitte, Forward Deployed Engineers (FDE) don’t just build AI... .... Work you’ll do As a Senior Frontier GenAI FDE, you will work side... ..., safety, latency, cost, and model risk.Deliver production-quality... ...with MLOps/LLMOps practices: evaluation frameworks, model monitoring,...Local areaVisa sponsorship- ...vast talent network trains frontier AI models in the same way teachers teach... ...the Role As a Software Engineer on the Automations team at... ...from MCP server design to evaluation frameworks — working directly... ...(or equivalent modern language). Experience building with...LanguageFull timeWork at officeImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Frontier Language Model Evaluation Engineer. Be the first to apply!
- language arts San Francisco, CA
- natural language processing San Francisco, CA
- language internship San Francisco, CA
- language subtitle San Francisco, CA
- dual language San Francisco, CA
- language San Francisco, CA
- language manager San Francisco, CA
- foreign language San Francisco, CA
- speech language San Francisco, CA
- language consultant San Francisco, CA



