Frontier Language Model Evaluation Engineer
Artificial Analysis, Inc.
Artificial Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets, scoring systems, and evaluation infrastructure applicable across major models released by labs worldwide. Based in San Francisco or flexible to other major cities, this role offers equity and the opportunity to shape how industry measures frontier AI capability while collaborating with leading researchers and #J-18808-Ljbffr Artificial Analysis, Inc.
- ...Description - Member Of Technical Staff (Language Model Evaluations) Location: San Francisco (preferred),... ...company. We support labs, engineers and enterprises to understand AI capabilities... ...of AI, they are actively shaping the frontier. Our benchmarks and analysis are...Language
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against...SuggestedPart timeImmediate start
$180k - $225k
...transformation through frontier AI systems that solve... ...ever. New foundation models, reasoning techniques,... ...remains one of the hardest engineering challenges.As a... ...customers to design, evaluate, and deploy intelligent... ...latest advances in large language models, reasoning, retrieval...LanguageFull time- ...are looking for strong engineers to join our team and... ...and benchmarking new models as they are released on... ...error modes of models, evaluate their strengths and weaknesses... ..., how to use large language models in practice.... ...pushing the frontier. Unicorn companies that...LanguageWork experience placementRelocationRelocation packageShift work
- ...applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research, design experiments and collaborate with research engineers, foundation-model developers and domain experts to turn open-ended questions...LanguageRelocation
- YO AI Labs is seeking a Turkish Bilingual Expert to support a language and AI training project. This contractor role is remote, with flexible hours to evaluate Turkish audio for nativeness, fluency, pronunciation, and intonation. Provide clear feedback in English, justify...LanguageRemote jobFor contractorsFlexible hours
$98k - $140k
...products. You’ll work with product and engineering teams to build systems to define... ...context engineering, designing evaluation systems, and analyzing data. This... ...of that you'll shape Notion's model strategy and work directly with frontier AI labs (OpenAI, Anthropic, Google...Live inLocal area$140k - $185k
...coverage and rapid sprints around model releases. About the Role... ...operate leaderboards that evaluate LLMs: test new model releases... ...relative strengths. Strong engineering fundamentals and a track record... ...experience benchmarking large language models or creating evaluation...LanguageFull timeRelocation package- A cutting-edge AI firm in San Francisco seeks a VLM Post-Training Owner to lead enterprise engagements and enhance vision-language models. The role combines project ownership with technical execution, ensuring quality data generation and customer satisfaction in AI solutions...Language
$136.44k - $265.11k
...data, and run AI agents and models directly in their workflows.... ...re a team focused on making frontier AI models better at science.... ....You’ll build the datasets, evaluations, and systems that help close... ...the intersection of software engineering, biology, and frontier AI:...Work at officeLocal areaMonday to FridayShift work- ...is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab... ...other standards, using your bilingual language skills. Ideal candidates have 3+ years as a music producer or audio engineer, a college degree in music, and...LanguagePart timeImmediate start10 hours per week
$220k - $320k
...net]( trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the... ...-to-end: distillation, training, evaluation, and planet-scale hosting. We... ...a well-funded ten-person team of engineers who work in-person in downtown San...LanguageWork at office- Mercor is hiring experienced musicians to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated music... ...by quality and originality, and evaluating natural language usage and regional expressions. #J-18808-Ljbffr ObsidianLanguage
- ...development of the most advanced AI models. 1. Overview We are hiring an... ...teams, improving how frontier AI models reason about real legal... ...challenging legal tasks and evaluation sets, and help build legal-... ...Hands-on working use of large language models in your professional...LanguageFull timeContract workPart timeFreelanceInternshipLive inRelocationRelocation package
- ...seeking a Dutch Bilingual Expert for a remote, contractor role. You will evaluate AI-generated Dutch speech, assess nativeness and accuracy, and provide detailed feedback to support language model improvements. No prior AI experience is required; strong Dutch and English...LanguageRemote jobFor contractorsFlexible hours
- ...standards, and research advanced defenses against emerging threats. Applicants should have deep expertise in LLM safety, strong software engineering skills, and relevant academic qualifications in AI or related fields. This position is pivotal for paving the way for safe AI...
$110k - $130k
...AI teams train and run models on the right data. Our... ..., annotates, and evaluates data across the full AI... ...of 100+ working at the frontier of AI and have raised... ...As a Customer Support Engineer, you will be working with... ...with SQL and scripting languages such as Python or...Language- ...RoleAs a Forward Deployed Engineer you will work closely... ...wants to be at the frontier of enterprise AI adoption... ...solutions.Evaluate build-vs-buy options for... ...using modern programming languages such as Python, JavaScript... ...platforms, enterprise data models, or regulated...LanguageFull timeTemporary workLocal areaImmediate start
- ...seeking a Hebrew Bilingual Expert to contribute to a language and AI training project focused on Hebrew audio evaluation. You will assess nativeness, fluency, and... ...linguistic quality, and provide feedback to help improve model understanding and generation. The role is remote...LanguageRemote jobContract work
- ...invites you to join a talent-dense team in San Francisco with 5.5M+ users and a rapidly growing platform. You will define how frontier AI models are measured, design new benchmarks, run experiments, and publish analyses that become industry gold standards. We sponsor...RelocationVisa sponsorship
$155.4k - $233.2k
...services operate. By combining frontier agentic AI, an enterprise-... ...bar behind Harvey’s human evaluations. As we scale globally, the volume... ...enough that Product, Engineering, and AI Research act on it to... ...quality across multiple markets, languages, or jurisdictions Built...LanguageContract work- ...bilingual experts to contribute to an AI training project. You will evaluate Polish audio content, assess nativeness, fluency, pronunciation... ...remotely, collaborate with coordinators, and help improve Polish-language understanding and generation. #J-18808-Ljbffr YO AI LabsRemote job
$315k
We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand... ...Safety Level (ASL) to assign to models. Research leads on this team... ...workstreams, we would value experience with language model agents, although this is not...LanguageCurrently hiringWork at officeImmediate startHome officeVisa sponsorshipRelocation package$218.5k - $288k
...Applied Scientist specializing in Small Language Models and AI Training, you will lead... ...You will work closely with research, engineering, and product teams to advance model training... ...language models.Design, implement, and evaluate model training experiments to improve...LanguageWork at officeFlexible hours3 days per week- YO AI Labs is seeking a Korean Language Expert contractor to contribute linguistic expertise for AI data projects. You will evaluate, translate, and annotate Korean content to help improve model understanding and natural language generation. This remote role welcomes candidates...LanguageRemote jobFor contractors
- ...Welo Data is seeking Data Labeling Associates in California to evaluate AI outputs and ensure cultural context and safety in Arabic datasets. This role requires professional-level proficiency in Portuguese (Brazil), a bachelor's degree, and at least 2 years of experience...
$35 per hour
Location : Remote Fluent Language Skills Required:... ...experts who probe AI models with adversarial inputs... ...development, reverse engineering Socio-technical risk:... ...customer AI systems Evaluation coverage expands: more... ...AI red teaming at the frontier of safety Play a direct...LanguageRemote job- Founding Engineer - AI-Native Fintech B2B Platform Harrison Clarke... .... The platform uses large language models, agentic workflows and real-... ...team grows Contribute to the evaluation and integration of LLMs,... ...founders High-impact work at the frontier of AI and financial...LanguageFlexible hours
- ...against real customer SKUs. Distill frontier vision‑language‑action models for edge deployment, retaining... ...Competencies And Skills We’re looking for engineers who combine strong robotics and ML... ...data pipelines, training systems, evaluation frameworks, observability tooling,...LanguageRelocation package
- ...our Sales, Product and Agent Engineering teams to both drive our GTM function... .... What You'll Bring ~ Language fluency - professional... ...intensity are compatible, and we model it in our actions and... ...job description. We strive to evaluate all applicants consistently without...LanguageFull timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Frontier Language Model Evaluation Engineer. Be the first to apply!
- natural language processing San Francisco, CA
- dual language San Francisco, CA
- language analyst San Francisco, CA
- language consultant San Francisco, CA
- foreign language San Francisco, CA
- language manager San Francisco, CA
- language subtitle San Francisco, CA
- language internship San Francisco, CA
- language arts San Francisco, CA
- language San Francisco, CA

