Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Model Evaluation Program Lead

$300k - $320k

Anthropic

About the role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be crucial in assessing the performance, capabilities, limitations, and potential risks of our AI models. Working closely with our Research, Trust & Safety, Frontier Redteaming, and Policy teams, you will drive high-priority evaluation projects to build new processes, align metrics with policy, and track measurable progress. You will help build and adapt the model evaluation program to ensure model deployments are rigorous and aligned with our commitment to responsible AI development. The ideal candidate will have a strong technical background and experience managing cross-functional programs in AI development, ML engineering, or related fields. You’ll be joining a team of Technical Program Managers who own and drive cross-functional programs that align to the company’s top priorities. In this role, you’ll have the opportunity to make a foundational impact as you contribute the scaling of a centralized TPM function for the company. Extremely strong soft skills are paramount, as our team is front and center in driving lots of company-wide changes and top priority initiatives that require generating buy-in, balancing various opinions, and competing for attention in our rapidly scaling environment. This role is a great fit for someone who has both seen excellence at scale and operated in rapidly scaling, high-ambiguity teams and scope. We are seeking candidates with deep TPM expertise but who are comfortable acting as adaptable generalists who add value fast. We excel at maintaining a broad view of our work but diving deep into the details when necessary. We understand business goals, translate and organize them into technical programs and projects, and drive execution. We are adept at engaging with both non-technical and technical stakeholders at all levels of the company, including executive leadership. In this role, you will have the opportunity to shape the development of advanced AI systems and contribute to Anthropic's mission of ensuring that AI benefits all of humanity. If you are passionate about responsible AI development, have a strong technical background, and thrive in a fast-paced, collaborative environment, we'd love to hear from you. Responsibilities: Partner with teams like Frontier Risk Evaluations, Security, and Trust & Safety to develop and implement comprehensive evaluation protocols for our latest frontier AI models Build a single source of truth for tracking all types of model evaluations as required by our Responsible Scaling Policy, AI safety institutes, the White House, and others Develop and maintain procedures for conducting evaluations, including designing test suites, coordinating red team exercises, and analyzing results Create and manage dashboards and reporting systems to track model performance, safety metrics, and evaluation outcomes across different AI systems and versions Lead cross-functional workshops to identify potential risks and edge cases for evaluation, ensuring thorough coverage of AI capabilities and limitations Coordinate with external partners and industry standards bodies to align our evaluation practices with emerging best practices in responsible AI development Provide detailed status reports, identifying technical risks, dependencies, and areas requiring additional support Facilitate communication and coordination between technical workstreams and stakeholders Continuously identify opportunities for technical process improvements and implement changes as needed Stay up-to-date with the latest developments in AI safety, ML engineering, and related fields to ensure the program remains at the forefront of responsible AI development You might be a good fit if you: Have several years of experience in technical program management, with a track record of successfully delivering complex technical programs, preferably in AI development, ML engineering, or related fields Have experience executing technical programs that require systems and engineering-level knowledge. Have exceptionally strong interpersonal and communication skills that enable you to influence without authority, build cross-organizational support, cooperation and action around initiatives and process adoption. Have experience prompt engineering on language models Have experience designing and/or running evaluations on Large Language Models Have knowledge of emerging AI governance frameworks and best practices Have a high threshold for navigating ambiguity and are able to balance setting strategic priorities with rapid, high-quality execution. Thrive in unstructured environments, and have a knack for bringing order to chaos. The expected salary range for this position is: Annual Salary:

$300,000—$320,000 USD

Logistics Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. US visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate; operations roles are especially difficult to support. But if we make you an offer, we will make every effort to get you into the United States, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Compensation and Benefits* Anthropic’s compensation package consists of three elements: salary, equity, and benefits. We are committed to pay fairness and aim for these three elements collectively to be highly competitive with market rates. Equity - For eligible roles, equity will be a major component of the total compensation. We aim to offer higher-than-average equity compensation for a company of our size, and communicate equity amounts at the time of offer issuance. US Benefits - The following benefits are for our US-based employees: Optional equity donation matching. Comprehensive health, dental, and vision insurance for you and all your dependents. 401(k) plan with 4% matching. 22 weeks of paid parental leave. Unlimited PTO – most staff take between 4-6 weeks each year, sometimes more! Stipends for education, home office improvements, commuting, and wellness. Fertility benefits via Carrot. Daily lunches and snacks in our office. Relocation support for those moving to the Bay Area. UK Benefits - The following benefits are for our UK-based employees: Optional equity donation matching. Private health, dental, and vision insurance for you and your dependents. Pension contribution (matching 4% of your salary). 21 weeks of paid parental leave. Unlimited PTO – most staff take between 4-6 weeks each year, sometimes more! Health cash plan. Life insurance and income protection. Daily lunches and snacks in our office. #J-18808-Ljbffr Anthropic

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Model Evaluation Program Lead in Seattle, WA vacancy
  • A local arts and culture organization is seeking an Evaluator to develop and implement an evaluation framework for measuring the impact of various programs. This role includes guiding the assessment report process and overseeing data collection methods. Successful candidates... 
    Suggested
    Local area

    4Culture - Seattle - Hybrid

    Seattle, WA
    2 days ago
  • Alignerr is seeking a Quantitative Analyst to evaluate and improve AI-generated mathematical outputs for finance applications. This fully remote...  ...AI reasons about risk and forecasting. You will analyze models for validity, assess performance, and validate data pipelines... 
    Suggested
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Alignerr

    Seattle, WA
    3 days ago
  • Welo Data is hiring a Data Labeling Associate in Washington with expertise in evaluating AI systems. This role focuses on providing structured feedback on model outputs and requires strong writing skills and attention to detail. As a full-time employee, you will engage... 
    Suggested
    Full time

    Welo Data

    Seattle, WA
    5 days ago
  •  ...for building machine learning models and systems to protect our...  ...models represents the frontier of AI today. This topic pioneers...  ...systems as a plus- Proficiency in programming languages such as Python,...  ....- Experience with evaluation of AI systems, LLM application... 
    Suggested
    Flexible hours
    Shift work

    TikTok

    Seattle, WA
    4 days ago
  • Welo Data is seeking a Data Labeling Associate in Washington State. This full-time role involves evaluating AI model outputs and improving data quality. The ideal candidate should have native-level Australian English proficiency, a bachelor’s degree, and strong analytical... 
    Suggested
    Full time

    Welo Data

    Seattle, WA
    5 days ago
  • $100 per hour

    A leading technology firm is seeking finance experts to enhance AI models. Responsibilities include evaluating performance in capital markets and creating assessment rubrics. Candidates should have 2+ years in finance fields like investment banking and possess strong financial... 
    Hourly pay
    Remote work
    10 hours per week

    Turing

    Seattle, WA
    12 hours ago
  • Oracle is seeking a seasoned strategist to lead and execute large cross-functional programs within a fast-moving, highly technical environment. You will drive the rhythm of business, align diverse teams toward a common north star, and remove blockers to maintain steady... 

    Oracle

    Seattle, WA
    5 days ago
  • $54 - $67 per hour

    A global player support organization is hiring a Knowledge Program Manager to lead knowledge management practices enhancing player and support experiences. The role requires expertise in AI and knowledge management with a focus on developing strategies for player-facing... 
    Hourly pay

    Onward Search

    Seattle, WA
    2 days ago
  • Google Inc. in Washington DC / Seattle, WA is seeking an experienced Manager, Responsible AI Strategic Programs to lead high-impact, cross-functional AI strategy initiatives. You will shape programs, drive outcomes, and influence stakeholders across Trust and Safety, product... 
    Flexible hours

    Google Inc.

    Seattle, WA
    3 days ago
  •  ...in Seattle seeks a senior researcher to develop novel AI testing methodologies and evaluation frameworks. You will partner with data science and engineering...  ...qualitative and quantitative inquiry, and requires leading safety initiatives with cross-functional teams across... 

    Google

    Seattle, WA
    2 days ago
  • $202.16k - $368.22k

     ...compilation technologies for AI foundation models. Responsibilities Design...  ...-scale model training, evaluation, and inference. Optimize distributed...  ...or change these benefits programs at any time, with or...  ...we have launched industry‑leading general foundation models and... 
    Temporary work
    Internship
    Local area

    ByteDance

    Seattle, WA
    2 days ago
  • $232.56k - $427.5k

     ...Design and build scalable infrastructure for large‑scale model training, evaluation, and inference. Optimize distributed training systems across...  ...reserves the right to modify or change these benefits programs at any time, with or without notice. Legal & EEO Qualified... 
    Temporary work
    Local area

    ByteDance

    Seattle, WA
    6 days ago
  • $202.16k - $368.22k

     ...Building a next-generation big model as a service platform to serve...  ...training and alignment, model evaluation, test-time scaling, agent...  ...foundation models and data-centric AI, particularly in how large...  ...reserves the right to modify these programs at any time, with or without... 
    Temporary work
    Internship
    Local area

    ByteDance

    Seattle, WA
    2 days ago
  •  ...full project life cycle, from up-front evaluation and planning studies through design and...  ...searching for an Industrial Wastewater (IWW) Program Lead - Oil, Gas, and Chemicals Practice....  ....Preferred Qualifications Process modeling experience.Existing relationships with... 
    For contractors

    HDR

    Seattle, WA
    2 days ago
  • $148.7k - $201.2k

     ...implementing AWS's global environmental programs — defining risk management...  ...program development, and lead cross-regional execution from...  ...teamAWS Infrastructure Services (AIS) owns the design, planning,...  ...assessments, environmental permit evaluations and applications, EIAs and... 
    For contractors
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $258.8k - $304.2k

     ...Transaction Services Advisory Lead, Software & AI West Monroe is seeking a...  ..., and C-suite executives to evaluate software companies and technology...  ..., and technology operating models while identifying...  ...our employee stock ownership program and be eligible to receive annual... 
    Work at office
    Local area
    Immediate start
    Flexible hours
    2 days per week

    West Monroe Partners

    Seattle, WA
    7 hours ago
  •  ...Vice President and Applied AI/ML Lead, you’ll play a pivotal role...  ...and iteration.Design scalable model pipelines for document ingestion...  ...text and documents.Define evaluation strategies and success...  ...outcomes in production.Strong programming skills in Python and experience... 

    JP Morgan Chase

    Seattle, WA
    2 days ago
  • $148.7k - $201.2k

     ...motivated Senior Manufacturing Environmental Program Integration Lead to support AWS Manufacturing, Test,...  ..., and action ownership.8. Performance Evaluation & Reporting - Establish and own a...  ...the teamAWS Infrastructure Services (AIS) owns the design, planning, delivery,... 
    Flexible hours
    Day shift

    Amazon

    Seattle, WA
    22 hours ago
  • $167.1k - $226.1k

    Build AI systems that help Amazon make better sustainability...  ...than applying an existing model: they require new...  ...test hypotheses, establish evaluation standards, and lead solutions from early experimentation...  ...can scale across multiple programs. This role shapes not just... 
    Flexible hours

    Amazon

    Seattle, WA
    22 hours ago
  • $172.5k - $260.1k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...career at the company leading workforce...  ...level security and access model in Tableau Next so every...  ...recruiters assess and evaluate candidates’ resumes and...  ...well including: time off programs, medical, dental, vision... 
    Full time

    Salesforce

    Seattle, WA
    2 days ago
  •  ...Senior Principal, Anthropic AI SolutionsAI Systems &...  ...Partner Solution Lead, you’ll partner with cross...  ...and acting as the SME on model usage, MCP integrations...  ...co-lead AI enablement programs: designing training...  ...integrations) and the ability to evaluate and guide... 
    Temporary work
    Work at office
    Local area

    Slalom

    Seattle, WA
    7 hours ago
  • $148.7k - $201.2k

     ...team of technical infrastructure program managers, software engineers,...  ....AWS Infrastructure Services (AIS) designs, delivers, and...  ...quickly and often involve long lead time inputs. You will partner...  ...Develop constrained-planning models that surface risks and levers... 
    Interim role
    Flexible hours

    Amazon

    Seattle, WA
    7 hours ago
  • $102.3k - $161.76k

     ...About the Role As a Staff AI FinOps Governance Lead, you will lead financial governance...  ..., and cross-functional program leadership to ensure AI...  ...value. Develop forecasting models and long-range financial plans...  ...utilization. Evaluate tradeoffs between reserved... 

    Ultimate Software

    Seattle, WA
    1 day ago
  • $142.8k - $147.9k

     ...challenges, application of AI technologies, and...  ...to adapt, innovate, and lead in an evolving landscape...  ...Production, Not Just Pilots Evaluate emerging AI platforms,...  ...CI/CD, monitoring, or model evaluation Sia invests...  ...student loan repayment programs Paid parental leave;... 
    Work at office
    Immediate start
    Worldwide
    3 days per week

    SIA

    Seattle, WA
    3 days ago
  •  ...is seeking a Student Researcher in Seattle to conduct research on infrastructure for AI foundation models. This role requires pursuing a PhD in computer science and strong programming skills, focusing on efficiency and reliability in large-scale systems. Interns enjoy... 
    Internship

    Pangleglobal

    Seattle, WA
    4 days ago
  •  ...strategic decisions focusing on three key areas: AI transformation, resource efficiency &...  ...level.- Ecosystem Diagnosis & Metrics: Evaluate and analyze the current state of the...  ...Hands-on experience in software development, programming, front-end, or back-end engineering is... 
    Shift work

    TikTok

    Seattle, WA
    4 days ago
  • $150.1k - $227k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...career at the company leading workforce...  ...design reviews, threat modeling, secure code reviews,...  ...potential attack vectors, evaluate business risk, and recommend...  ...in one or more programming or scripting languages... 
    Full time

    Salesforce

    Seattle, WA
    2 days ago
  •  ...committed to being a positive role model for the youth we serve....  ...for our 2026-2027 After-School Programming. The 2026-2027 school year...  ...developmentServe as the staff lead for assigned programs at designated...  ...Act.Applicants will be evaluated on the basis of education and... 
    Temporary work
    Local area
    Flexible hours
    Shift work

    City of Burien

    Seattle, WA
    6 days ago
  • $233.6k - $362.2k

     ...The Portfolio and Platform Lead will play a central role in managing...  ...the lifecycle of MNCNH’s AI projects—from proof‑of‑concept...  ...maternal, newborn, and child health programs. The individual will bridge...  ...of AI product development, evaluation, and adoption, including the product... 
    Relocation

    Bill & Melinda Gates Foundation

    Seattle, WA
    12 hours ago
  • $171k - $248k

     ...years of experience in AI testing or research, data...  ...and scale testing programs, streamlining the launch...  ...designing first-of-their-kind evaluations for Google’s most...  ...assessing novel foundational model capabilities as they...  ...In this role, you will lead the development of... 
    Temporary work

    Google

    Seattle, WA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Model Evaluation Program Lead. Be the first to apply!