Member of Technical Staff (Language Model Evaluations)
Artificial Analysis
Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don't just measure the cutting edge of AI, they are actively shaping the frontier. Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist. We are a team of 40+, on track to triple by year end, backed by Nat Friedman (Github, Meta), Daniel Gross (SSI), Andrew Ng (Google Brain, DeepLearning.ai, Amazon), Adam D'Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders. The Opportunity Language model evaluation is the sharpest question in AI: what can these systems actually do? Our answers, from the Artificial Analysis Intelligence Index to AA-Omniscience, AA-Briefcase and our coding agent evaluations, are the reference the industry uses. We’re hiring Members of Technical Staff to build the next generation of them. This is a role for people who want to build frontier benchmarks: designing evaluations that stay ahead of frontier capabilities, constructing datasets that resist contamination, and measuring what everyone else has not yet worked out how to measure. You will run your work across every major model as it releases and publish results the whole industry reads. The center of the role is building. Analysis and lab collaboration wrap around the evaluation work, with our commercial team owning client relationships day to day. What You’ll Do Design Next-Generation Frontier Evals: Conceive and ship the next generation of frontier evaluations, like AA-Briefcase and AA-Omniscience, across reasoning, knowledge, coding, agentic capability and beyond Build Evaluation Datasets and Infrastructure: Construct the datasets, harnesses and scoring systems behind our benchmarks, engineered for contamination resistance and repeatability at frontier scale Shape the Future Intelligence Index: The evaluations you build will contribute to future versions of the Artificial Analysis Intelligence Index and other areas of our platform, defining how the industry measures frontier capability Publish Influential Analysis: Produce the reports, indexes and data visualizations that shape how the industry understands language model progress Work with Frontier Labs on Pre-Release Models: Benchmark the leading labs’ systems, including pre-release and newly launched models, working directly with their research teams; our commercial team owns client relationships day to day, so your time stays on the science Evaluate Every Major Model: Run our evaluation suite across frontier releases as they land, and own the integrity of the results the industry quotes Become AI-Native: Embrace an AI-native workflow, using cutting-edge AI tools to generate leverage in a fast-changing industry and maintain our competitive edge in AI benchmarking What We’re Looking For You have deep, hands-on experience evaluating language models and strong opinions about why most benchmarks fail. Backgrounds include: evaluation and benchmarking teams at AI labs; research or engineering roles at evaluation-focused organizations; ML engineers who have built evaluation harnesses and datasets in production; or academic researchers in NLP and ML evaluation with a strong record of published work. Required: 3+ years of relevant professional experience, across industry or research Strong analytical and critical thinking skills Strong Python, with hands-on experience running evaluation harnesses and building datasets Deep familiarity with the LLM evaluation landscape: the major benchmarks and their failure modes, contamination, preference-based methods, and agentic evaluation Strong statistical grounding: you know when a result is signal and when it is noise Genuine, demonstrable interest and knowledge of Frontier AI. We want people who have informed opinions about where AI is heading, not just people who use AI tools Why Artificial Analysis? Shape how AI gets built: The leading AI labs track our benchmarks and use them to guide their development priorities. Your work will directly influence the direction of AI. Become a world expert in AI: You will evaluate every major model, across every major capability, as they are released. Very few roles offer this breadth of exposure to frontier AI. Work with the most important players in AI: You’ll manage relationships with teams at the leading AI labs and major enterprises as a trusted, independent voice. Join at a defining moment: We’re 40+ people, on track to double by end of year, backed by some of the most connected investors in AI. The people who join now will shape the product, the team, and the strategy as we scale. Competitive compensation including equity Interested in this role? #J-18808-Ljbffr Artificial Analysis
- ...AI research company on a search for a Member of Technical Staff focused on AI Safet y. The company... ...next-generation open-weight foundation models with a mission to make advanced AI... ...help define how advanced AI systems are evaluated, stress-tested, and safely deployed....Suggested
- ...Member of Technical Staff @ Lotus AI Who we are Lotus AI is a groundbreaking... ...role. You may work across model training and fine-tuning,... ...high-quality training and evaluation datasets from clinical interactions... ...speech-to-text, natural language understanding, and text-to-...Language
- ...industry leaders. The Opportunity Language model inference is the fastest-moving market... ...reference point, and we’re hiring a Member of Technical Staff to drive them. You’ll own coverage... ...Become a world expert in AI: You will evaluate every major model, across every...LanguageShift work
- ...trust continually improving models. This includes trajectory visualization, evaluation workflows, monitoring dashboards... ...will help define the visual language, interaction patterns, and... ...building our team of founding Members of Technical Staff to design the frontier of continually...Language
- ...rigor of distributed systems, model architecture, and numerics... .... About the Role As a Member of Technical Staff, Mechanistic Interpretability... ...study how multimodal genome language models represent, process,... ...learning systems, improved model evaluations, and ultimately, mastery...LanguageLocal area
$180k
Member of Technical Staff - Virology As molecular data generation and frontier model intelligence grows, new approaches to data analysis are... ...and biologists to build evaluations that measure whether AI systems... ...used or evaluated Large Language Models, biological foundation...LanguageFull timeWork at officeFlexible hoursNight shift- ...rigor of distributed systems, model architecture, and numerics... .... About the Role As a Member of Technical Staff focused on Protein Structure... ...model development, training, evaluation, and scientific analysis, with... ...and fine‑tune protein language models, geometric neural networks...LanguageLocal area
- Member of Technical Staff, Applied AI The opportunity We are looking for a Member... ...expertise in generative modelling to work at the interface... ...technical concepts into clear language for scientific and non-... ...unique data challenges, evaluation paradigms and scientific workflows...LanguageFlexible hours
- ...cutting‑edge foundation AI models and end‑to‑end products that... ...our new flagship vision‑language model: Consistently outperforms... ..., and join the team. As a member of technical staff with a focus on multimodal... ...have experience building evaluations to measure their performance...LanguageFull timeWork at officeLocal areaRemote workHome office
$150k - $300k
...from frontier agentic models to the infra that enables... ...our ability to evaluate and serve models trained... ...training stack. Core Technical Responsibilities LLM Serving... ...Systems Performance Languages: Rust, C++. Data &... ...development and encourage team members to contribute to the...LanguageWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- ...robotic intelligence. As a Member of Technical Staff, you'll be at the forefront... ...developing breakthrough foundation models that enable robots to... ..., end-to-end vision-language-action models, efficient model... ...infrastructure to prototype and evaluate at scale Collaborate with...LanguageLocal area
- Artificial Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets, scoring systems, and evaluation infrastructure applicable across major models released...LanguageWorldwideFlexible hours
- ...encodings, foundation models, ML systems, fine-tuning... ...& images Test, evaluate, and characterize natural language AI systems Requires proficiency... ...We're a small, all-technical team, all working at... ...gets the same title: Member of Technical Staff. Compensation varies with...LanguagePermanent employmentWork at officeVisa sponsorshipWork visa
$250k
Eragon — Member of Technical Staff Type: Full-time | On-site | San Francisco, CA... ...It post-trains open-source models on a customer's own data, integrates... ...action through natural language — all deployed in the... ...Model development: Fine-tune, evaluate, and work with ML models in...LanguageFull timeH1bWork at officeLocal areaVisa sponsorship- ...and backers at . About the Role As a Member of Technical Staff, you will help invent and build the... ...cloud operations. Design, implement, and evaluate algorithms using large‑scale... ...developing large software systems in languages such as C++, Python, Go, or Rust. Experience...LanguageWork from homeFlexible hours2 days per week
- ...frontier of interactive world models: systems that generate,... ...intelligent systems to learn, evaluate, and interact within... ...AI The Role We are hiring a Member of Technical Staff to lead reinforcement learning... ...team to improve large vision-language and code-generating agents...Language
$200k - $400k
...field of AI-based simulation, proving it is possible to model human behavior with high accuracy. Today, we are... ...Adam D'Angelo, and Guillermo Rauch. About the Role As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement systems that...Flexible hours- ...We're hiring software engineers for the Model Behavior team to help shape how Perplexity... ...engineering fundamentals, and a technical understanding of LLM-driven and agentic... ...external partners. Experience designing evaluations or benchmarks for AI systems. #J-18808-...
- Job You will own the training pipeline behind the models that power both Parallel’s search stack and Parallel’s agents. On the search... ...real product usage to high‑quality training data, fine‑tune and evaluate these models rigorously, and ship them safely to traffic used...Work at officeVisa sponsorship
$200k - $400k
...5x, built a new foundation model for human behavior that has... ...chance. About the Role As a Member of Technical Staff (MTS) in Research, you will... ...across the stack to train, evaluate, deploy, and monitor our... ...at the forefront of modern language model training and fine-tuning...LanguageFlexible hours- ...core engineering team to build and train models that understand these systems, optimize... ...calling to make all decisions on a network. Evaluate model performance over real‑world... ...and software stack. Want to build the technical DNA of a new applied research org from the...
$150k - $250k
...Product Hunt), Charlie Songhurst (Board Member, Meta), and Michael Jones (Former Chair... ...Develop custom performance and quality evaluations for our agents Minimum... ...development (Python) # Experience integrating language models into AI applications Preferred Qualifications...LanguageFull timeInternshipWorldwide$256k - $276k
Member of Technical Staff, AI Agent Development Lead Join to apply for the Member of Technical Staff... ...leveraging state‑of‑the‑art language models and associated technologies. Collaborate... ...environment of collaboration and innovation. Evaluate new tools, frameworks, and...LanguageFull timeFlexible hours$150k - $350k
...Output has built a biological reasoning model that understands biology at the scale and... ...scalable computational tools and methods that evaluate synthetic feasibility across generated... ...chemistry terms, translating between the language of generative AI and the language of...Language- ...It's the product. Why the title is Member of Technical Staff Because they'd rather hire the person... ...We don't care what it said. # Your language or stack. We have opinions and we'll... .... # An AI or ML background. We run models against live money and we'd like you to...LanguageWork at officeRemote workVisa sponsorshipFlexible hours
$203.5k - $299.3k
...org's culture and operating model — how we publish, how we collaborate... ...runs, and large-scale evaluation sweeps Full research infrastructure... ...— vision, speech, and language applied to merchant catalogs,... ...economies shapes how our team members move quickly, learn, and reiterate...LanguageHourly payWork at officeLocal areaRemote workFlexible hours- ...from the methods that shape how models reason to the systems that... ...a training run, the next on evaluation infrastructure, and the next... ...across modalities, including language, structured and scientific data... ...deployment and iteration. Act as a technical resource within the team,...LanguageLive in
$200k - $300k
...high agency. The Role Nimble is looking for a Member of Technical Staff to help us advance our robotics moonshot by designing,... ...developing, training, and implementing robotic foundation models, Vision-Language-Action Models, general-purpose robotic AI models, and...LanguageLocal areaImmediate startFlexible hoursWeekend work- ...The Opportunity Liquid AI's models ship inside real products,... ...and turn them into trained, evaluated, production-ready model checkpoints... ...: You can explain something technical you built from first... ...reliable across all supported languages. Generate, clean, and analyze...LanguageFull timeInternshipShift work
- ...and build on. We build open models that let anyone control... ...model capabilities. As a member of the Data Team, your mission... ...crawlers, and solving the unique technical challenges of collecting... ...is used in training and evaluating large language models Experience with distributed...LanguageWork at officeVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff (Language Model Evaluations). Be the first to apply!
- work from home technical support specialist San Francisco, CA
- product support technician San Francisco, CA
- helpdesk support technician San Francisco, CA
- help desk assistant San Francisco, CA
- senior technical associate San Francisco, CA
- IT help desk technician San Francisco, CA
- technical solutions specialist San Francisco, CA
- desktop support analyst San Francisco, CA
- trade support analyst San Francisco, CA
- senior IT support technician San Francisco, CA



