Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Frontier Language Model Evaluation Engineer

Artificial Analysis, Inc.

Artificial Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets, scoring systems, and evaluation infrastructure applicable across major models released by labs worldwide. Based in San Francisco or flexible to other major cities, this role offers equity and the opportunity to shape how industry measures frontier AI capability while collaborating with leading researchers and #J-18808-Ljbffr Artificial Analysis, Inc.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Frontier Language Model Evaluation Engineer in San Francisco, CA vacancy
  •  ...benchmarking company. We support labs, engineers and enterprises to understand AI...  ...of AI, they are actively shaping the frontier. Our benchmarks and analysis are trusted...  ...other industry leaders. The Opportunity Language model evaluation is the sharpest question in AI: what... 
    Language

    Artificial Analysis

    San Francisco, CA
    1 day ago
  • $180k - $225k

     ...transformation through frontier AI systems that solve...  ...ever. New foundation models, reasoning techniques,...  ...remains one of the hardest engineering challenges.As a...  ...customers to design, evaluate, and deploy intelligent...  ...latest advances in large language models, reasoning, retrieval... 
    Language
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $98k - $140k

     ...products. You’ll work with product and engineering teams to build systems to define...  ...context engineering, designing evaluation systems, and analyzing data. This...  ...of that you'll shape Notion's model strategy and work directly with frontier AI labs (OpenAI, Anthropic, Google... 
    Suggested
    Live in
    Local area

    Notion Labs

    San Francisco, CA
    1 day ago
  •  ...Technical Staff to build the next generation of AI benchmarks. You will design frontier evaluations, construct evaluation datasets, and publish analyses that set industry standards for model capability and reliability. You will work across leading labs and enterprises,... 
    Suggested

    Artificial Analysis

    San Francisco, CA
    1 day ago
  •  ...remotely. The role focuses on fact-checking and generating evaluation data, requiring native fluency in Urdu and strong English writing...  ...a bachelor’s degree and significant experience with large language models. Responsibilities also include independently assessing... 
    Language
    Remote job

    Mercor

    San Francisco, CA
    1 day ago
  • $238k - $302k

     ...simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language Models (LLMs) and Vision-Language...  ...are looking for quantitatively-minded engineers to research and propose new ways to... 
    Language
    Full time
    Remote work

    Waymo

    San Francisco, CA
    18 hours ago
  • $15 - $20 per hour

     ...tools . Generate high-quality human evaluation data by identifying response strengths,...  ...and completeness of responses. Ensure model responses align with expected conversational...  .... Significant experience using large language models (LLMs). Excellent writing... 
    Language
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    15 days ago
  • Mercor is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab....  ...of music in your bilingual language. Ideal candidates have 3+ years in music production or audio engineering, a degree in music, and fluency in French... 
    Language
    Part time
    Immediate start
    10 hours per week

    Mercor

    San Francisco, CA
    4 days ago
  • A cutting-edge AI firm in San Francisco seeks a VLM Post-Training Owner to lead enterprise engagements and enhance vision-language models. The role combines project ownership with technical execution, ensuring quality data generation and customer satisfaction in AI solutions... 
    Language

    Liquid AI

    San Francisco, CA
    3 days ago
  • Mercor is seeking experienced musicians to evaluate generative musical AI models in collaboration with a leading AI lab. You will assess model outputs across categories of music in your bilingual language, focusing on lyrics, voice generation, and audio quality. The role... 
    Language
    Part time

    Mercor Inc

    San Francisco, CA
    3 days ago
  • Welo Data in San Francisco seeks a full-time AI Evaluator with professional proficiency in Portuguese (Portugal) and experience in Generative AI safety. The role involves critiquing AI outputs, identifying biases, and refining evaluation frameworks. Candidates should possess... 
    Language
    Full time

    Welo Data

    San Francisco, CA
    4 days ago
  •  ...is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab...  ...other standards, using your bilingual language skills. Ideal candidates have 3+ years as a music producer or audio engineer, a college degree in music, and... 
    Language
    Part time
    Immediate start
    10 hours per week

    Mercor Inc

    San Francisco, CA
    4 days ago
  • $220k - $320k

     ...net]( trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the...  ...-to-end: distillation, training, evaluation, and planet-scale hosting. We...  ...a well-funded ten-person team of engineers who work in-person in downtown San... 
    Language
    Full time
    Work at office

    Inference

    San Francisco, CA
    3 days ago
  •  ...is seeking experienced musicians to evaluate generative musical AI models in collaboration with a leading AI...  ...categories of music in your bilingual language and contribute to structured...  ...years in music production or audio engineering, a music degree, and native or near... 
    Language
    Part time
    Immediate start
    10 hours per week

    Mercor

    San Francisco, CA
    4 days ago
  •  ...standards, and research advanced defenses against emerging threats. Applicants should have deep expertise in LLM safety, strong software engineering skills, and relevant academic qualifications in AI or related fields. This position is pivotal for paving the way for safe AI... 

    Xcede

    San Francisco, CA
    2 days ago
  • $110k - $130k

     ...AI teams train and run models on the right data. Our...  ..., annotates, and evaluates data across the full AI...  ...of 100+ working at the frontier of AI and have raised...  ...As a Customer Support Engineer, you will be working with...  ...with SQL and scripting languages such as Python or... 
    Language

    Encord

    San Francisco, CA
    3 days ago
  •  ...RoleAs a Forward Deployed Engineer you will work closely...  ...wants to be at the frontier of enterprise AI adoption...  ...solutions.Evaluate build-vs-buy options for...  ...using modern programming languages such as Python, JavaScript...  ...platforms, enterprise data models, or regulated... 
    Language
    Full time
    Temporary work
    Local area
    Immediate start

    Innovaccer

    San Francisco, CA
    3 days ago
  • $155.4k - $233.2k

     ...services operate. By combining frontier agentic AI, an enterprise-...  ...bar behind Harvey's human evaluations. As we scale globally, the volume...  ...enough that Product, Engineering, and AI Research act on it to...  ...quality across multiple markets, languages, or jurisdictions Built... 
    Language
    Contract work

    Neura Market

    San Francisco, CA
    2 days ago
  • $80 per hour

     ...and Jack Dorsey . Position: Data Engineer (Coding Agent Experience) Type: Contract...  ...Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations involving... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    1 day ago
  • $218.5k - $288k

     ...Applied Scientist specializing in Small Language Models and AI Training, you will lead...  ...You will work closely with research, engineering, and product teams to advance model training...  ...language models.Design, implement, and evaluate model training experiments to improve... 
    Language
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    1 day ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand...  ...Safety Level (ASL) to assign to models. Research leads on this team...  ...workstreams, we would value experience with language model agents, although this is not... 
    Language
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    San Francisco, CA
    18 hours ago
  • Engineering Manager, Foundation Model Inference (FMAPI)RDQ427R519At Databricks, we are driven by a passion to...  ...to serve, scale, and optimize frontier models with enterprise-grade reliability...  ..., gender identity or expression, language, national origin, physical and mental... 
    Language
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $120k - $200k

     ...The Role As a Founding Engineer, you will work...  ...and precedent Create evaluation infrastructure for high...  ...system architecture and model behavior to UI, deployment...  ...deeply curious about frontier multimodal and agentic...  ...infrastructure, and modern language and vision models. How... 
    Language
    For contractors
    Local area

    Red Cedar Ventures

    San Francisco, CA
    2 days ago
  • Obsidian is seeking a Spanish Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project. You will handle...  ..., and evaluation tasks to help train and benchmark advanced language models. The ideal candidate should have strong writing skills,... 
    Language
    Part time
    10 hours per week

    Obsidian

    San Francisco, CA
    2 days ago
  •  ...software and foundation models enable vehicles to...  ...autonomy. Working at the frontier of a still-undefined...  ...collaborate with research and engineering teams to define what “...  ...pretraining and evaluation Manage and mentor a...  ...datasets involving video, language, lidar, radar and... 
    Language
    Full time
    Work at office
    Remote work
    Work from home
    Visa sponsorship
    Relocation package
    Flexible hours

    GrabJobs

    San Francisco, CA
    2 days ago
  • $300k - $320k

     ...Program Manager to lead our AI model evaluation initiatives across multiple...  ...Research, Trust & Safety, Frontier Redteaming, and Policy...  ...programs in AI development, ML engineering, or related fields. You’ll...  ...prompt engineering on language models Have experience designing... 
    Language
    Work at office
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    San Francisco, CA
    1 day ago
  • $134.5k - $265.1k

     ...At Deloitte, Forward Deployed Engineers (FDE) don’t just build AI...  ...0/30/2026.Work you’ll doAs a Frontier GenAI FDE, you will work side...  ..., safety, latency, cost, and model risk.Deliver production-quality...  ...with MLOps/LLMOps practices: evaluation frameworks, model monitoring,... 
    Local area

    Deloitte

    San Francisco, CA
    18 hours ago
  • Welo Data is seeking Data Labeling Associates in San Francisco to evaluate AI systems focused on Arabic language nuances. The role includes model evaluation, identifying biases in datasets, and ensuring quality across global AI workflows. Candidates should have native French... 

    Welo Data

    San Francisco, CA
    18 hours ago
  • $155.6k - $306.8k

     ...At Deloitte, Forward Deployed Engineers (FDE) don’t just build AI...  .... Work you’ll do As a Senior Frontier GenAI FDE, you will work side...  ..., safety, latency, cost, and model risk.Deliver production-quality...  ...with MLOps/LLMOps practices: evaluation frameworks, model monitoring,... 
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    18 hours ago
  •  ...vast talent network trains frontier AI models in the same way teachers teach...  ...the Role As a Software Engineer on the Automations team at...  ...from MCP server design to evaluation frameworks — working directly...  ...(or equivalent modern language). Experience building with... 
    Language
    Full time
    Work at office
    Immediate start

    Mercor

    San Francisco, CA
    18 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Frontier Language Model Evaluation Engineer. Be the first to apply!