Member of Technical Staff, Evaluation Execution
$328k - $402kMETR
Job Description
Job Description
About METR
We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
What this role looks likeRunning models on tasks. Often this means integrating models into our agent scaffolds, running them on our infrastructure and checking the results carefully. (METR both develops our own tasks internally and runs external evaluations.)
Communicating results and takeaways. This includes designing useful graphs, writing up conclusions for different audiences (system cards, risk reports, regulators, X, etc), and having great takes on what matters for risk.
Building software to improve our evaluations. We don't just try and run the same evaluation over and over again. We also run faster, more informative evaluations over time; this means making the right investments (with the support of our platform team).
Project management. Live evaluations require keeping track of a bunch of threads and staying organized. With our recent risk report process, we were running many evaluations at once.
Strong and professional communication. We run important and sensitive evaluations, and so the team needs to coordinate with METR leadership, lab contacts, regulators, and others.
As part of informing the world about risk from frontier AI systems, METR often runs and publishes evaluations of frontier models.
Our evaluations are a central tool the world uses to understand AI progress. Our Time Horizon methodology has been included in system cards, called an "obsession" by the NYT, has wide reach online, and is used by governments to inform national policy.
We’re expanding the ambition and scale of our evaluations. We have recently begun to measure model propensities and monitorability, and we are increasing the speed, reliability, and quantity of evaluations we aim to do so that we can keep the world informed.
Time Horizon is close to saturation, so we’re currently working on Time Horizon 2.0 , which we expect to be running on models over the next 6 to 18 months.
We’re gearing up for our first large-scale publication on monitorability, which we believe will be similar to TH in helping folks understand trends over time.
We spent the past three months working on a large, industry-wide third-party risk assessment program - which includes us collecting information (and running evaluations!) for both monitorability and propensities/alignment. We expect to do much more work as part of our own risk assessment programs in the future.
In general, many ambitious impact stories for METR require us having the capacity to run many more evaluations than we have run historically. For example, while our evaluations currently inform many key decisionmakers about AI capabilities, they are not yet consistently run with the scale, reliability, and speed necessary to play concrete, codified roles in regulatory frameworks. Unlocking this capacity is part of the near-future vision for evaluation execution.
Required skillsSoftware engineering. You're a strong engineer with solid infra fundamentals. You can dig into unfamiliar systems, debug from logs, and identify and fix performance bottlenecks.
Speed and scrappiness. You get things done quickly. You’re able to quickly identify what 80/20 looks like, and then do that.
High attention to detail. You read closely, can spot bugs in transcripts, and pay attention to the important fiddly bits.
Research understanding and taste. You understand research ideas and priorities, and have good intuitions for which plots are informative and which analyses are worth running to poke at the data.
Strong external communicator. You communicate well with external stakeholders, and we trust you to stay on the ball with communications with, e.g., lab contacts.
Project management. You can juggle many balls at once, keep stakeholders updated, and track and anticipate blockers.
Strong writing ability. You can be a solid contributor to METR’s writeups of evaluation results, see e.g. our GPT-5 report.
We're looking for both more junior/mid-level and more senior versions of the role. For junior/mid-level, the compensation ranges from $328k-402k, and for more senior, the compensation ranges from $402k-578k.
For very experienced and exceptional candidates, we are open to exploring paying much higher than this stated range.
The listed range applies to the base salary for this role. METR also has a host of benefits:
- The office: Catered lunch and dinner daily; in-office gym and shower
- Relocation support: Stipend for moving to the Bay Area
- Time-off and leave: Unlimited PTO and 21-week parental leave for new parents
- Commuter benefit: Monthly transit/parking stipend and an annual Uber budget
- Professional development benefit: for training, courses, conferences, and AI safety education
- Mental health benefit: for therapy, medication, and other mental health expenses
- Wellness benefit: for gym memberships and other wellness expenses
- Work equipment benefit: for home office and workstation equipment expenses
Our Culture
METR is a mission-driven organization. We believe our work can meaningfully shape humanity's future for the better, and we want to be the best people in the world doing this work. We have a tight-knit, collaborative research culture rooted in truth-seeking and integrity. We're fiercely committed to producing high-quality, trustworthy science. We're honest and transparent about our results, especially when they may go against the grain. We've earned trust as reliable partners who handle confidential information with care. We maintain a low-ego, drama-free environment focused on what matters.
Hybrid Preferred: Our technical team members are in our office in Berkeley 3-5 days/week. We would ideally like for you to be in person too, but we are happy to be flexible here. If you lack US work authorization and would like to work in-person, we can likely sponsor a cap-exempt H-1B visa for this role.
We encourage you to apply even if your background may not seem like the perfect fit! We would rather review a larger pool of applications than risk missing out on a promising candidate for the position.
We are committed to diversity and equal opportunity in all aspects of our hiring process. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We welcome and encourage all qualified candidates to apply for our open positions.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
$285.55k
...our time. What We're Looking For The Evaluation Execution team at METR focuses on productionizing... ..., scalable systems and make sound technical decisions. You lead large projects from... ...Hybrid Requirements Our technical team members are in our office in Berkeley 3‑5 days...SuggestedH1bWork at officeWork from homeHome officeRelocation package3 days per week$140k - $200k
...Member Of Technical Staff San Francisco Bay Area Shape The Future Of Ai At Labelbox, we're... ...environment rewards high agency and rapid execution. Continuous Growth: Every role... ...AI agents rely on during training and evaluation, the terminals, browsers, and tool-...SuggestedImmediate startFlexible hours$227.5k - $401k
...Member of Technical Staff San Francisco This is Adyen Adyen provides payments, data, and financial... ...Innovate and Deploy: Drive the execution of Adyen's AI strategy, focusing on... ...Multi-step Reasoning (DABStep), which evaluates AI agents on real-world data analysis...SuggestedWork at officeImmediate startRelocationFlexible hours$130k - $220k
...to help them hire. Title of Role: Member of Technical Staff (AI Benchmarking) Location: San Francisco... ...of the most important independent evaluators of frontier AI systems. The... ...What You Will Do Design and execute AI benchmarking and evaluation projects...SuggestedWork at officeRemote workVisa sponsorshipRelocation package- ...individuals who tackle unique technical challenges at scale and... ...technology sector. As a Member of Technical Staff , you will operate with a... ...Innovate and Deploy: Drive the execution of Adyen\'s AI strategy ,... ...(DABStep), which evaluates AI agents on real-world data...SuggestedFlexible hours
- ...work will define what cutting edge means. We're hiring Members of Technical Staff to design the evaluations that set the standard for how AI is measured,... ...Benchmarking Product Development: Structure, design and execute projects to evaluate AI systems and technologies, including...
- ...-growth technology company to hire a Member of Technical Staff (MTS) in Toronto. This is a senior, Staff... ...data requirements. Workflow & Execution Systems: Build reliable execution infrastructure... ...-driven processes. Observability & Evaluation: Develop tracing, metrics, structured...
$160k - $220k
...Job Description Job Description Member of Technical Staff Company: Bluejay Location: San... ...build systems that simulate, analyze and evaluate conversational AI agents across voice... ...and the ability to lead technical execution ~ An exceptional signal: a strong bachelor...Full timeWork at office$104.6k - $154.5k
...students, faculty, and staff, who are among the... ...outpacing our ability to evaluate it, eroding the rigor... ...scope, and deliver on technical requirements that require... .... They are integral members of a diverse team, co-... ...literature search, analysis execution and review, and...Full timeTraineeshipWork at office- ...builds the simulation, observability, and evaluation infrastructure to make voice the... ...available. About the Role: As a Senior Member of Technical Staff, you'll play a pivotal technical role... ..., can lead a project's technical execution end to end, and have built high-...Full timeWork at officeVisa sponsorship
$225k - $300k
...Description Job Description Senior Member of Technical Staff Company: Bluejay Location: San... ...systems that simulate, analyze and evaluate conversational AI agents across voice... ...safe and aligned. Lead technical execution of projects end to end. Work on-site...Full timeWork at office$220k - $350k
...Microsoft. The Role We’re hiring a Member of Technical Staff – AI/ML to design, build, and deploy... ...insights, recommend next steps, and execute approved tasks — not just chatbots... ...feature engineering, model selection, evaluation, calibration Have strong opinions...Full timeWorldwideFlexible hours- ...looking for engineers who want to build systems where AI is the execution layer and humans design, guide, and govern those systems.... ...to: inspect system state debug agent behavior evaluate outcomes steer system direction Implement feedback...
$180k - $250k
...providers rely on benchmarks.bio to evaluate and improve their AI systems... ...where engineers talk through technical challenges and stay aligned... ...will own a product area and execute, working closely with... ...Nathan • Founders & Chief of Staff: Alfredo, Kyle, Kenny, Jordan...Work at officeVisa sponsorshipWork visa- ...Member of Technical Staff @ Lotus AI Who we are Lotus AI is a groundbreaking primary care app that integrates your medical records, AI... ...curation pipelines that produce high-quality training and evaluation datasets from clinical interactions. Voice and Video...
- ...The role As a Member of Technical Staff at Sainapse, you will build the shared platform that turns enterprise customer workflows into reliable... ...models, integration and tool frameworks, permissions, evaluations, and deployment primitives. Build reliable backend and...Full timeVisa sponsorship
$200k
...You can refer people through our form. We're hiring a member of technical staff to work closely with the founding team. You'll shape both... ...learning: Contribute to agent science research, and create evaluations for agent performance and behavior Frontend engineering...Immediate start- ...dive in headfirst into new problems. About the Role As a Member of Technical Staff at Phonic, you'll work across the full range of what it takes... ...ambiguous to shipped without needing a fully-scoped spec. Execution speed: you ship, learn, and iterate quickly rather than...Work at officeShift work
$150k - $300k
...help us build and deploy our agents at scale. We’re a small, execution‑focused team. If you're excited by fast feedback loops, messy... ...outcomes with our design partners — your work will hit production Evaluate and integrate emerging AI frameworks, tools, and best...Work at office$200k - $300k
...fast-growing AI inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems that make LLM inference... ...there, you decide what to prove, build it, win the technical evaluation on the customer's own workload, and keep it running in...H1bWork at office$150k - $220k
# Founding Member of Technical Staff (MTS)Bay Area, CAFull-time$150k-$220k + equity## About UsVizopsAI is the secure runtime for custom enterprise... ...power continuous optimization loops for AI agents—from evaluation pipelines and data/trace infrastructure to APIs that...- ...thought partner. Backed by Y Combinator. About the Role As a Member of Technical Staff, you will build the core infrastructure and environments... ...the full stack, from designing reward functions and evaluation pipelines to building the tooling that lets us scale environment...
- ...Bedrock to start Bluejay. Bluejay builds evaluation infrastructure for Voice AI. In just... ...with 2+ YOE who want to play a pivotal technical role in Bluejay’s team. We want senior,... ...details, can lead a project’s technical execution, and have built high-traffic features at...Work at office
$2,000 per month
...faster than anyone in the world. Job Summary We’re hiring Members of Technical Staff — the general research/engineering seat. If you have a... ...Responsibilities Run research bets on Build data: train, evaluate, ablate, and ship what works Design and train models on in...Work at officeRelocation package$402.05k
...to several weeks at a time, likely alongside 1-4 other METR staff. Between exercises, you'll practice, develop the general methodology... ...environment focused on what matters. Hybrid Preferred: Our technical team members are in our office in Berkeley 3-5 days/week. We would...H1bWork at officeWork from homeHome officeRelocation packageFlexible hours3 days per week- ...uses Shapes every single day, and everyone talks to users. Member of Technical Staff is the title we use for engineers who own hard problems... ...React surfaces — feed, chat, profiles, and rooms Fine-tuning, evaluating, and routing across frontier models like Opus 4.7 and GPT...
- San Francisco, United States | Posted on 09/16/2026 Member of Technical Staff - AI & Core Engineering About the Opportunity FDE Team builds and... ..., retrieval systems, and cloud infrastructure. Build evaluation, monitoring, testing, and quality frameworks for production...Full time
$200k - $300.09k
...The Role The OpenClaw Foundation is seeking exceptional Members of Technical Staff (MTS) to serve as full‑time maintainers, builders, and stewards... ...exploring the human‑agent relationship Prototype and evaluate new capabilities Translate research findings into production...Full time- ...resolution no single person can. We're looking for founding members of technical staff: engineers who will co-own Hivemind's technical foundation... ...rather own a large ambiguous problem end to end than execute a spec someone else wrote. You've shipped and scaled robust...Full time
- ...About the Role As a Member of Technical Staff at Orpex, you'll work on agentic infrastructure and agent-building systems that power the next... ...implement scalable infrastructure for agent orchestration and execution. Build frameworks and tools that enable teams to...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff, Evaluation Execution. Be the first to apply!
- tech aide Berkeley, CA
- IT help desk technician Berkeley, CA
- customer support analyst Berkeley, CA
- work from home technical support specialist Berkeley, CA
- mri tech aide Berkeley, CA
- support technician Berkeley, CA
- support analyst Berkeley, CA
- desktop support analyst Berkeley, CA
- help desk technical support Berkeley, CA
- technical analyst Berkeley, CA


