Member of Technical Staff, Evaluation Execution
$328k - $402kMETR
Job Description
Job Description
[Due to capacity constraints, we may be slow to get back to you about this role.] About METR
We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
What this role looks likeRunning models on tasks. Often this means integrating models into our agent scaffolds, running them on our infrastructure and checking the results carefully. (METR both develops our own tasks internally and runs external evaluations.)
Communicating results and takeaways. This includes designing useful graphs, writing up conclusions for different audiences (system cards, risk reports, regulators, X, etc), and having great takes on what matters for risk.
Building software to improve our evaluations. We don't just try and run the same evaluation over and over again. We also run faster, more informative evaluations over time; this means making the right investments (with the support of our platform team).
Project management. Live evaluations require keeping track of a bunch of threads and staying organized. With our recent risk report process, we were running many evaluations at once.
Strong and professional communication. We run important and sensitive evaluations, and so the team needs to coordinate with METR leadership, lab contacts, regulators, and others.
As part of informing the world about risk from frontier AI systems, METR often runs and publishes evaluations of frontier models.
Our evaluations are a central tool the world uses to understand AI progress. Our Time Horizon methodology has been included in system cards, called an "obsession" by the NYT, has wide reach online, and is used by governments to inform national policy.
We’re expanding the ambition and scale of our evaluations. We have recently begun to measure model propensities and monitorability, and we are increasing the speed, reliability, and quantity of evaluations we aim to do so that we can keep the world informed.
Time Horizon is close to saturation, so we’re currently working on Time Horizon 2.0 , which we expect to be running on models over the next 6 to 18 months.
We’re gearing up for our first large-scale publication on monitorability, which we believe will be similar to TH in helping folks understand trends over time.
We spent the past three months working on a large, industry-wide third-party risk assessment program - which includes us collecting information (and running evaluations!) for both monitorability and propensities/alignment. We expect to do much more work as part of our own risk assessment programs in the future.
In general, many ambitious impact stories for METR require us having the capacity to run many more evaluations than we have run historically. For example, while our evaluations currently inform many key decisionmakers about AI capabilities, they are not yet consistently run with the scale, reliability, and speed necessary to play concrete, codified roles in regulatory frameworks. Unlocking this capacity is part of the near-future vision for evaluation execution.
Required skillsSoftware engineering. You're a strong engineer with solid infra fundamentals. You can dig into unfamiliar systems, debug from logs, and identify and fix performance bottlenecks.
Speed and scrappiness. You get things done quickly. You’re able to quickly identify what 80/20 looks like, and then do that.
High attention to detail. You read closely, can spot bugs in transcripts, and pay attention to the important fiddly bits.
Research understanding and taste. You understand research ideas and priorities, and have good intuitions for which plots are informative and which analyses are worth running to poke at the data.
Strong external communicator. You communicate well with external stakeholders, and we trust you to stay on the ball with communications with, e.g., lab contacts.
Project management. You can juggle many balls at once, keep stakeholders updated, and track and anticipate blockers.
Strong writing ability. You can be a solid contributor to METR’s writeups of evaluation results, see e.g. our GPT-5 report.
We're looking for both more junior/mid-level and more senior versions of the role. For junior/mid-level, the compensation ranges from $328k-402k, and for more senior, the compensation ranges from $402k-578k.
For very experienced and exceptional candidates, we are open to exploring paying much higher than this stated range.
METR also has a host of benefits:
The office: Catered lunch and dinner daily; in-office gym and shower
Relocation support: Stipend for moving to the Bay Area
Time-off and leave: Unlimited PTO and 21-week parental leave for new parents
Commuter benefit: Monthly transit/parking stipend and an annual Uber budget
Professional development benefit: for training, courses, conferences, and AI safety education
Wellness benefit: for gym memberships, exercise equipment, therapy, medication, and other wellness expenses
Work equipment benefit: for home office and workstation equipment expenses
Our Culture
METR is a mission-driven organization. We believe our work can meaningfully shape humanity's future for the better, and we want to be the best people in the world doing this work. We have a tight-knit, collaborative research culture rooted in truth-seeking and integrity. We're fiercely committed to producing high-quality, trustworthy science. We're honest and transparent about our results, especially when they may go against the grain. We've earned trust as reliable partners who handle confidential information with care. We maintain a low-ego, drama-free environment focused on what matters.
Hybrid Preferred: Our technical team members are in our office in Berkeley 3-5 days/week. We would ideally like for you to be in person too, but we are happy to be flexible here. If you lack US work authorization and would like to work in-person, we can likely sponsor a cap-exempt H-1B visa for this role.
We encourage you to apply even if your background may not seem like the perfect fit! We would rather review a larger pool of applications than risk missing out on a promising candidate for the position.
We are committed to diversity and equal opportunity in all aspects of our hiring process. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We welcome and encourage all qualified candidates to apply for our open positions.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
$227.5k - $401k
...Member of Technical Staff San Francisco This is Adyen Adyen provides payments, data, and financial... ...Innovate and Deploy: Drive the execution of Adyen's AI strategy, focusing on... ...Multi-step Reasoning (DABStep), which evaluates AI agents on real-world data analysis...SuggestedWork at officeImmediate startRelocationFlexible hours- ...backgrounds may be a fit. THE ROLE As a Member of Technical Staff, you may work on projects that require strong execution, communication, analytical judgment, and the... ...profile context that helps recruiting teams evaluate fit. WHO THIS IS FOR This is a...Suggested
$130k - $220k
...to help them hire. Title of Role: Member of Technical Staff (AI Benchmarking) Location: San Francisco... ...of the most important independent evaluators of frontier AI systems. The... ...What You Will Do Design and execute AI benchmarking and evaluation projects...SuggestedWork at officeRemote workVisa sponsorshipRelocation package$140k - $200k
...Member Of Technical Staff San Francisco Bay Area Shape The Future Of Ai At Labelbox, we're... ...environment rewards high agency and rapid execution. Continuous Growth: Every role... ...AI agents rely on during training and evaluation, the terminals, browsers, and tool-...SuggestedImmediate startFlexible hours- ...will define what cutting edge means. We're hiring Members of Technical Staff to design the evaluations that set the standard for how AI is measured,... ...Benchmarking Product Development: Structure, design and execute projects to evaluate AI systems and technologies, including...Suggested
$160k - $220k
...Job Description Job Description Member of Technical Staff Company: Bluejay Location: San... ...build systems that simulate, analyze and evaluate conversational AI agents across voice... ...and the ability to lead technical execution ~ An exceptional signal: a strong bachelor...Full timeWork at office$225k - $300k
...Description Job Description Senior Member of Technical Staff Company: Bluejay Location: San... ...systems that simulate, analyze and evaluate conversational AI agents across voice... ...safe and aligned. Lead technical execution of projects end to end. Work on-site...Full timeWork at office- ...builds the simulation, observability, and evaluation infrastructure to make voice the... ...available. About the Role: As a Senior Member of Technical Staff, you'll play a pivotal technical role... ..., can lead a project's technical execution end to end, and have built high-...Full timeWork at officeVisa sponsorship
$180k - $250k
...providers rely on benchmarks.bio to evaluate and improve their AI systems... ...where engineers talk through technical challenges and stay aligned... ...will own a product area and execute, working closely with... ...Nathan • Founders & Chief of Staff: Alfredo, Kyle, Kenny, Jordan...Work at officeVisa sponsorshipWork visa- ...uses Shapes every single day, and everyone talks to users. Member of Technical Staff is the title we use for engineers who own hard problems... ...surfaces — feed, chat, profiles, and rooms Fine-tuning, evaluating, and routing across frontier models like Opus 4.7 and GPT...
- ...Member of Technical Staff @ Lotus AI Who we are Lotus AI is a groundbreaking primary care app that integrates your medical records, AI... ...curation pipelines that produce high-quality training and evaluation datasets from clinical interactions. Voice and Video...
- ...San Francisco, United States | Posted on 09/16/2026 Member of Technical Staff – AI & Core Engineering About the Opportunity FDE... ..., retrieval systems, and cloud infrastructure. Build evaluation, monitoring, testing, and quality frameworks for production...Full time
$200k
...You can refer people through our form. We're hiring a member of technical staff to work closely with the founding team. You'll shape both... ...learning: Contribute to agent science research, and create evaluations for agent performance and behavior Frontend engineering...Immediate start$150k - $300k
...business that is exploding as AI traverses the physical economy. What You'll Do As a Member of Technical Staff on our Frontier Data team, you’ll build the environments, evaluations, and datasets that expand what frontier models can actually do. This is work at the...Work at office$2,000 per month
...than anyone in the world. Job Summary We’re hiring Members of Technical Staff — the general research/engineering seat. If you have a seriously... ...Run research bets on Build data: train, evaluate, ablate, and ship what works Design and train models on...Work at officeRelocation package- ...The role As a Member of Technical Staff at Sainapse, you will build the shared platform that turns enterprise customer workflows into reliable... ...models, integration and tool frameworks, permissions, evaluations, and deployment primitives. Build reliable backend and...Full timeVisa sponsorship
- ...looking for engineers who want to build systems where AI is the execution layer and humans design, guide, and govern those systems.... ...to: inspect system state debug agent behavior evaluate outcomes steer system direction Implement feedback...
$150k - $220k
# Founding Member of Technical Staff (MTS)Bay Area, CAFull-time$150k-$220k + equity## About UsVizopsAI is the secure runtime for custom enterprise... ...power continuous optimization loops for AI agents—from evaluation pipelines and data/trace infrastructure to APIs that...- ...in headfirst into new problems. About the Role As a Member of Technical Staff at Phonic, you'll work across the full range of what it takes... ...to shipped without needing a fully-scoped spec. Execution speed: you ship, learn, and iterate quickly rather than overplanning...Work at officeShift work
- ...Our client, an AI healthcare software startup, is seeking a Member of Technical Staff (Research Scientist) to join their team. In this role, you'll be at the forefront of developing benchmarks and evaluation methodologies for large language models, helping shape how the...H1bRelocation
- ...products, and build the sustainable systems to execute accordingly. Cost efficiency is a key... ...good work. Great interpersonal and technical communication. Please don\'t use LLMs to... ...Share an online whiteboard with a team member and work through a technical problem. We...Work at office
- ...About the Role As a Member of Technical Staff at Orpex, you'll work on agentic infrastructure and agent-building systems that power the next... ...implement scalable infrastructure for agent orchestration and execution. Build frameworks and tools that enable teams to...
$200k - $350k
...directly with frontier labs on their most consequential data, evaluation, and post-training challenges, building the systems that... ...useful for advancing AI. The Role We are hiring a Member of Technical Staff, Evals to help define how frontier AI systems are measured...Full timeWork at officeFlexible hours- ...to several weeks at a time, likely alongside 1-4 other METR staff. Between exercises, you'll practice, develop the general methodology... ...focused on what matters. Hybrid Requirements: Our technical team members are in our office in Berkeley 3-5 days/week. Please let us...H1bWork at officeWork from homeRelocationHome officeRelocation package3 days per week
- ...to several weeks at a time, likely alongside 1-4 other METR staff. Between exercises, you'll practice, develop the general methodology... ...focused on what matters. Hybrid Requirements: Our technical team members are in our office in Berkeley 3-5 days/week. Please let us...H1bWork at officeWork from homeHome officeRelocation package3 days per week
- ...Member Of Technical Staff, Machine Learning Drug discovery is a prediction problem. Scientists design molecules that they predict will be... ...machine learning fundamentals, deep learning architectures, and evaluation approaches You are proficient in standard Python-based...
$200k
...main tools for pacing the frontier safely. The role Members of Technical Staff own a research project end to end. You would reason carefully... ...hypotheses and then into large-scale experiments that you execute. Depending on your expertise you may focus more on...Full timeRemote workRelocationVisa sponsorshipWork visaRelocation package- ...Member of Technical Staff — DevOps We look for infrastructure engineers who are obsessed with engineering velocity. Everything we do — research... ...We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains....
- ...isn't a quality bar. It's the product. Why the title is Member of Technical Staff Because they'd rather hire the person than fill the box.... ...and that isn't flexible. You want a defined roadmap to execute against. We often don't know what the right thing to build...Work at officeRemote workVisa sponsorshipFlexible hours
$180k - $220k
...domain logic Design application specific workflows for compound evaluation, program prioritization, and multi modal evidence integration... ..., and hard work. We solve hard problems through focused daily execution Speed: We ship fast (2x/week) and improve continuously...Work at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff, Evaluation Execution. Be the first to apply!


