Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff (Language Model Evaluations)

Artificial Analysis

Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don't just measure the cutting edge of AI, they are actively shaping the frontier. Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist. We are a team of 40+, on track to triple by year end, backed by Nat Friedman (Github, Meta), Daniel Gross (SSI), Andrew Ng (Google Brain, DeepLearning.ai, Amazon), Adam D'Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders. The Opportunity Language model evaluation is the sharpest question in AI: what can these systems actually do? Our answers, from the Artificial Analysis Intelligence Index to AA-Omniscience, AA-Briefcase and our coding agent evaluations, are the reference the industry uses. We’re hiring Members of Technical Staff to build the next generation of them. This is a role for people who want to build frontier benchmarks: designing evaluations that stay ahead of frontier capabilities, constructing datasets that resist contamination, and measuring what everyone else has not yet worked out how to measure. You will run your work across every major model as it releases and publish results the whole industry reads. The center of the role is building. Analysis and lab collaboration wrap around the evaluation work, with our commercial team owning client relationships day to day. What You’ll Do Design Next-Generation Frontier Evals: Conceive and ship the next generation of frontier evaluations, like AA-Briefcase and AA-Omniscience, across reasoning, knowledge, coding, agentic capability and beyond Build Evaluation Datasets and Infrastructure: Construct the datasets, harnesses and scoring systems behind our benchmarks, engineered for contamination resistance and repeatability at frontier scale Shape the Future Intelligence Index: The evaluations you build will contribute to future versions of the Artificial Analysis Intelligence Index and other areas of our platform, defining how the industry measures frontier capability Publish Influential Analysis: Produce the reports, indexes and data visualizations that shape how the industry understands language model progress Work with Frontier Labs on Pre-Release Models: Benchmark the leading labs’ systems, including pre-release and newly launched models, working directly with their research teams; our commercial team owns client relationships day to day, so your time stays on the science Evaluate Every Major Model: Run our evaluation suite across frontier releases as they land, and own the integrity of the results the industry quotes Become AI-Native: Embrace an AI-native workflow, using cutting-edge AI tools to generate leverage in a fast-changing industry and maintain our competitive edge in AI benchmarking What We’re Looking For You have deep, hands-on experience evaluating language models and strong opinions about why most benchmarks fail. Backgrounds include: evaluation and benchmarking teams at AI labs; research or engineering roles at evaluation-focused organizations; ML engineers who have built evaluation harnesses and datasets in production; or academic researchers in NLP and ML evaluation with a strong record of published work. Required: 3+ years of relevant professional experience, across industry or research Strong analytical and critical thinking skills Strong Python, with hands-on experience running evaluation harnesses and building datasets Deep familiarity with the LLM evaluation landscape: the major benchmarks and their failure modes, contamination, preference-based methods, and agentic evaluation Strong statistical grounding: you know when a result is signal and when it is noise Genuine, demonstrable interest and knowledge of Frontier AI. We want people who have informed opinions about where AI is heading, not just people who use AI tools Why Artificial Analysis? Shape how AI gets built: The leading AI labs track our benchmarks and use them to guide their development priorities. Your work will directly influence the direction of AI. Become a world expert in AI: You will evaluate every major model, across every major capability, as they are released. Very few roles offer this breadth of exposure to frontier AI. Work with the most important players in AI: You’ll manage relationships with teams at the leading AI labs and major enterprises as a trusted, independent voice. Join at a defining moment: We’re 40+ people, on track to double by end of year, backed by some of the most connected investors in AI. The people who join now will shape the product, the team, and the strategy as we scale. Competitive compensation including equity Interested in this role? #J-18808-Ljbffr Artificial Analysis

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff (Language Model Evaluations) in San Francisco, CA vacancy
  •  ...AI research company on a search for a Member of Technical Staff focused on AI Safet y. The company...  ...next-generation open-weight foundation models with a mission to make advanced AI...  ...help define how advanced AI systems are evaluated, stress-tested, and safely deployed.... 
    Suggested

    Xcede

    San Francisco, CA
    2 days ago
  •  ...Member of Technical Staff @ Lotus AI Who we are Lotus AI is a groundbreaking...  ...role. You may work across model training and fine-tuning,...  ...high-quality training and evaluation datasets from clinical interactions...  ...speech-to-text, natural language understanding, and text-to-... 
    Language

    Lotus Health AI, Inc

    San Francisco, CA
    3 days ago
  •  ...industry leaders. The Opportunity Language model inference is the fastest-moving market...  ...reference point, and we’re hiring a Member of Technical Staff to drive them. You’ll own coverage...  ...Become a world expert in AI: You will evaluate every major model, across every... 
    Language
    Shift work

    Artificial Analysis

    San Francisco, CA
    2 days ago
  •  ...trust continually improving models. This includes trajectory visualization, evaluation workflows, monitoring dashboards...  ...will help define the visual language, interaction patterns, and...  ...building our team of founding Members of Technical Staff to design the frontier of continually... 
    Language

    Trajectory

    San Francisco, CA
    2 days ago
  •  ...rigor of distributed systems, model architecture, and numerics...  .... About the Role As a Member of Technical Staff, Mechanistic Interpretability...  ...study how multimodal genome language models represent, process,...  ...learning systems, improved model evaluations, and ultimately, mastery... 
    Language
    Local area

    Radical Numerics

    San Francisco, CA
    3 days ago
  • $180k

    Member of Technical Staff - Virology As molecular data generation and frontier model intelligence grows, new approaches to data analysis are...  ...and biologists to build evaluations that measure whether AI systems...  ...used or evaluated Large Language Models, biological foundation... 
    Language
    Full time
    Work at office
    Flexible hours
    Night shift

    LatchBio

    San Francisco, CA
    3 days ago
  •  ...rigor of distributed systems, model architecture, and numerics...  .... About the Role As a Member of Technical Staff focused on Protein Structure...  ...model development, training, evaluation, and scientific analysis, with...  ...and fine‑tune protein language models, geometric neural networks... 
    Language
    Local area

    Radical Numerics

    San Francisco, CA
    2 days ago
  • Member of Technical Staff, Applied AI The opportunity We are looking for a Member...  ...expertise in generative modelling to work at the interface...  ...technical concepts into clear language for scientific and non-...  ...unique data challenges, evaluation paradigms and scientific workflows... 
    Language
    Flexible hours

    Latent Labs

    San Francisco, CA
    4 days ago
  •  ...cutting‑edge foundation AI models and end‑to‑end products that...  ...our new flagship vision‑language model: Consistently outperforms...  ..., and join the team. As a member of technical staff with a focus on multimodal...  ...have experience building evaluations to measure their performance... 
    Language
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    San Francisco, CA
    4 days ago
  • $150k - $300k

     ...from frontier agentic models to the infra that enables...  ...our ability to evaluate and serve models trained...  ...training stack. Core Technical Responsibilities LLM Serving...  ...Systems Performance Languages: Rust, C++. Data &...  ...development and encourage team members to contribute to the... 
    Language
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    3 days ago
  •  ...robotic intelligence. As a Member of Technical Staff, you'll be at the forefront...  ...developing breakthrough foundation models that enable robots to...  ..., end-to-end vision-language-action models, efficient model...  ...infrastructure to prototype and evaluate at scale Collaborate with... 
    Language
    Local area

    Amazon Science

    San Francisco, CA
    1 day ago
  • Artificial Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets, scoring systems, and evaluation infrastructure applicable across major models released... 
    Language
    Worldwide
    Flexible hours

    Artificial Analysis, Inc.

    San Francisco, CA
    4 days ago
  •  ...encodings, foundation models, ML systems, fine-tuning...  ...& images Test, evaluate, and characterize natural language AI systems Requires proficiency...  ...We're a small, all-technical team, all working at...  ...gets the same title: Member of Technical Staff. Compensation varies with... 
    Language
    Permanent employment
    Work at office
    Visa sponsorship
    Work visa

    Generation Alpha Transistor

    San Francisco, CA
    2 days ago
  • $250k

    Eragon — Member of Technical Staff Type: Full-time | On-site | San Francisco, CA...  ...It post-trains open-source models on a customer's own data, integrates...  ...action through natural language — all deployed in the...  ...Model development: Fine-tune, evaluate, and work with ML models in... 
    Language
    Full time
    H1b
    Work at office
    Local area
    Visa sponsorship

    davidjoseph-co

    San Francisco, CA
    3 days ago
  •  ...and backers at . About the Role As a Member of Technical Staff, you will help invent and build the...  ...cloud operations. Design, implement, and evaluate algorithms using large‑scale...  ...developing large software systems in languages such as C++, Python, Go, or Rust. Experience... 
    Language
    Work from home
    Flexible hours
    2 days per week

    Emerald AI

    San Francisco, CA
    1 day ago
  •  ...frontier of interactive world models: systems that generate,...  ...intelligent systems to learn, evaluate, and interact within...  ...AI The Role We are hiring a Member of Technical Staff to lead reinforcement learning...  ...team to improve large vision-language and code-generating agents... 
    Language

    Moonlake AI

    San Francisco, CA
    3 days ago
  • $200k - $400k

     ...field of AI-based simulation, proving it is possible to model human behavior with high accuracy. Today, we are...  ...Adam D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement systems that... 
    Flexible hours

    Simile

    San Francisco, CA
    6 hours ago
  •  ...We're hiring software engineers for the Model Behavior team to help shape how Perplexity...  ...engineering fundamentals, and a technical understanding of LLM-driven and agentic...  ...external partners. Experience designing evaluations or benchmarks for AI systems. #J-18808-... 

    Doist

    San Francisco, CA
    3 days ago
  • Job You will own the training pipeline behind the models that power both Parallel’s search stack and Parallel’s agents. On the search...  ...real product usage to high‑quality training data, fine‑tune and evaluate these models rigorously, and ship them safely to traffic used... 
    Work at office
    Visa sponsorship

    Parallel Web Systems

    San Francisco, CA
    4 days ago
  • $200k - $400k

     ...5x, built a new foundation model for human behavior that has...  ...chance. About the Role As a Member of Technical Staff (MTS) in Research, you will...  ...across the stack to train, evaluate, deploy, and monitor our...  ...at the forefront of modern language model training and fine-tuning... 
    Language
    Flexible hours

    Simile

    San Francisco, CA
    3 days ago
  •  ...core engineering team to build and train models that understand these systems, optimize...  ...calling to make all decisions on a network. Evaluate model performance over real‑world...  ...and software stack. Want to build the technical DNA of a new applied research org from the... 

    Meter

    San Francisco, CA
    3 days ago
  • $150k - $250k

     ...Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former Chair...  ...Develop custom performance and quality evaluations for our agents Minimum...  ...development (Python) # Experience integrating language models into AI applications Preferred Qualifications... 
    Language
    Full time
    Internship
    Worldwide

    Krew

    San Francisco, CA
    more than 2 months ago
  • $256k - $276k

    Member of Technical Staff, AI Agent Development Lead Join to apply for the Member of Technical Staff...  ...leveraging state‑of‑the‑art language models and associated technologies. Collaborate...  ...environment of collaboration and innovation. Evaluate new tools, frameworks, and... 
    Language
    Full time
    Flexible hours

    Postman

    San Francisco, CA
    4 days ago
  • $150k - $350k

     ...Output has built a biological reasoning model that understands biology at the scale and...  ...scalable computational tools and methods that evaluate synthetic feasibility across generated...  ...chemistry terms, translating between the language of generative AI and the language of... 
    Language

    Output Biosciences

    San Francisco, CA
    18 days ago
  •  ...It's the product. Why the title is Member of Technical Staff Because they'd rather hire the person...  ...We don't care what it said. # Your language or stack. We have opinions and we'll...  .... # An AI or ML background. We run models against live money and we'd like you to... 
    Language
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours

    EQL Tech

    San Francisco, CA
    16 days ago
  • $203.5k - $299.3k

     ...org's culture and operating model — how we publish, how we collaborate...  ...runs, and large-scale evaluation sweeps Full research infrastructure...  ...— vision, speech, and language applied to merchant catalogs,...  ...economies shapes how our team members move quickly, learn, and reiterate... 
    Language
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordashusa

    San Francisco, CA
    3 days ago
  •  ...from the methods that shape how models reason to the systems that...  ...a training run, the next on evaluation infrastructure, and the next...  ...across modalities, including language, structured and scientific data...  ...deployment and iteration. Act as a technical resource within the team,... 
    Language
    Live in

    Autopoiesis Sciences

    San Francisco, CA
    16 hours ago
  • $200k - $300k

     ...high agency. The Role Nimble is looking for a Member of Technical Staff to help us advance our robotics moonshot by designing,...  ...developing, training, and implementing robotic foundation models, Vision-Language-Action Models, general-purpose robotic AI models, and... 
    Language
    Local area
    Immediate start
    Flexible hours
    Weekend work

    Nimble Robotics

    San Francisco, CA
    2 days ago
  •  ...The Opportunity Liquid AI's models ship inside real products,...  ...and turn them into trained, evaluated, production-ready model checkpoints...  ...: You can explain something technical you built from first...  ...reliable across all supported languages. Generate, clean, and analyze... 
    Language
    Full time
    Internship
    Shift work

    Liquid Ai Inc

    San Francisco, CA
    16 hours ago
  •  ...and build on. We build open models that let anyone control...  ...model capabilities. As a member of the Data Team, your mission...  ...crawlers, and solving the unique technical challenges of collecting...  ...is used in training and evaluating large language models Experience with distributed... 
    Language
    Work at office
    Visa sponsorship

    Socket.dev

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff (Language Model Evaluations). Be the first to apply!