Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Model Evaluation Program Lead

$300k - $320k

Anthropic

About the role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be crucial in assessing the performance, capabilities, limitations, and potential risks of our AI models. Working closely with our Research, Trust & Safety, Frontier Redteaming, and Policy teams, you will drive high-priority evaluation projects to build new processes, align metrics with policy, and track measurable progress. You will help build and adapt the model evaluation program to ensure model deployments are rigorous and aligned with our commitment to responsible AI development. The ideal candidate will have a strong technical background and experience managing cross-functional programs in AI development, ML engineering, or related fields. You’ll be joining a team of Technical Program Managers who own and drive cross-functional programs that align to the company’s top priorities. In this role, you’ll have the opportunity to make a foundational impact as you contribute the scaling of a centralized TPM function for the company. Extremely strong soft skills are paramount, as our team is front and center in driving lots of company-wide changes and top priority initiatives that require generating buy-in, balancing various opinions, and competing for attention in our rapidly scaling environment. This role is a great fit for someone who has both seen excellence at scale and operated in rapidly scaling, high-ambiguity teams and scope. We are seeking candidates with deep TPM expertise but who are comfortable acting as adaptable generalists who add value fast. We excel at maintaining a broad view of our work but diving deep into the details when necessary. We understand business goals, translate and organize them into technical programs and projects, and drive execution. We are adept at engaging with both non-technical and technical stakeholders at all levels of the company, including executive leadership. In this role, you will have the opportunity to shape the development of advanced AI systems and contribute to Anthropic's mission of ensuring that AI benefits all of humanity. If you are passionate about responsible AI development, have a strong technical background, and thrive in a fast-paced, collaborative environment, we'd love to hear from you. Responsibilities: Partner with teams like Frontier Risk Evaluations, Security, and Trust & Safety to develop and implement comprehensive evaluation protocols for our latest frontier AI models Build a single source of truth for tracking all types of model evaluations as required by our Responsible Scaling Policy, AI safety institutes, the White House, and others Develop and maintain procedures for conducting evaluations, including designing test suites, coordinating red team exercises, and analyzing results Create and manage dashboards and reporting systems to track model performance, safety metrics, and evaluation outcomes across different AI systems and versions Lead cross-functional workshops to identify potential risks and edge cases for evaluation, ensuring thorough coverage of AI capabilities and limitations Coordinate with external partners and industry standards bodies to align our evaluation practices with emerging best practices in responsible AI development Provide detailed status reports, identifying technical risks, dependencies, and areas requiring additional support Facilitate communication and coordination between technical workstreams and stakeholders Continuously identify opportunities for technical process improvements and implement changes as needed Stay up-to-date with the latest developments in AI safety, ML engineering, and related fields to ensure the program remains at the forefront of responsible AI development You might be a good fit if you: Have several years of experience in technical program management, with a track record of successfully delivering complex technical programs, preferably in AI development, ML engineering, or related fields Have experience executing technical programs that require systems and engineering-level knowledge. Have exceptionally strong interpersonal and communication skills that enable you to influence without authority, build cross-organizational support, cooperation and action around initiatives and process adoption. Have experience prompt engineering on language models Have experience designing and/or running evaluations on Large Language Models Have knowledge of emerging AI governance frameworks and best practices Have a high threshold for navigating ambiguity and are able to balance setting strategic priorities with rapid, high-quality execution. Thrive in unstructured environments, and have a knack for bringing order to chaos. The expected salary range for this position is: Annual Salary:

$300,000—$320,000 USD

Logistics Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. US visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate; operations roles are especially difficult to support. But if we make you an offer, we will make every effort to get you into the United States, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Compensation and Benefits* Anthropic’s compensation package consists of three elements: salary, equity, and benefits. We are committed to pay fairness and aim for these three elements collectively to be highly competitive with market rates. Equity - For eligible roles, equity will be a major component of the total compensation. We aim to offer higher-than-average equity compensation for a company of our size, and communicate equity amounts at the time of offer issuance. US Benefits - The following benefits are for our US-based employees: Optional equity donation matching. Comprehensive health, dental, and vision insurance for you and all your dependents. 401(k) plan with 4% matching. 22 weeks of paid parental leave. Unlimited PTO – most staff take between 4-6 weeks each year, sometimes more! Stipends for education, home office improvements, commuting, and wellness. Fertility benefits via Carrot. Daily lunches and snacks in our office. Relocation support for those moving to the Bay Area. UK Benefits - The following benefits are for our UK-based employees: Optional equity donation matching. Private health, dental, and vision insurance for you and your dependents. Pension contribution (matching 4% of your salary). 21 weeks of paid parental leave. Unlimited PTO – most staff take between 4-6 weeks each year, sometimes more! Health cash plan. Life insurance and income protection. Daily lunches and snacks in our office. #J-18808-Ljbffr Anthropic

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Model Evaluation Program Lead in San Francisco, CA vacancy
  • $207k - $285k

    OpenAI is seeking a Technical Program Manager in San Francisco to lead initiatives that ensure the safety and robustness of its AI models. The role involves collaborating with diverse teams to turn risks into actionable plans. Ideal candidates will have experience in technical... 
    Suggested

    OpenAI

    San Francisco, CA
    2 days ago
  • Accenture is looking for a strategy and design professional to lead enterprise operating model initiatives in the United States. You will partner with...  ...ways of working and to guide end-to-end transformation programs. The role emphasizes organization design, client... 
    Suggested

    Accenture

    San Francisco, CA
    3 days ago
  •  ...professional to act as the primary liaison for model labs and partner ecosystems. You will...  ...infrastructure development, co-selling, and program management with a direct impact on revenue...  ..., and a proven track record with AI-related GTM strategies. #J-18808-Ljbffr Workman... 
    Suggested
    Contract work

    Workman Labs

    San Francisco, CA
    2 days ago
  • $238k - $302k

     ...across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in...  ...: ~ Proficiency in programming in Python or C++ ~ Experience with...  .... JAX, Tensorflow) Experience leading a team of Engineers The expected... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    20 hours ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Suggested
    Weekday work

    Mercor Inc

    San Francisco, CA
    20 hours ago
  •  ...in San Francisco is seeking a Research Program Manager who embeds with the Pre-training...  ...ML and Data teams to accelerate frontier model development. You will work alongside researchers...  ...data pipelines, experiments, and evaluation cycles. You bring 7+ years of technical... 

    CONFIDENTIAL Scovai

    San Francisco, CA
    2 days ago
  • $50 - $75 per hour

    A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in... 
    Hourly pay
    Contract work

    Mercor

    San Francisco, CA
    4 days ago
  • Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM... 
    Weekday work

    Mercor

    San Francisco, CA
    1 day ago
  • $350k

     ...Our first goal is to democratize frontier AI R&D across scientific disciplines. We...  ...AI research company and training our own models end-to-end. Our work spans areas such as...  ...looking for a research engineer to build the evaluation infrastructure that tells us whether our... 

    Mirendil

    San Francisco, CA
    4 days ago
  • Obsidian is partnering with an AI research initiative to support a Frontier Code...  ...project in San Francisco. You will evaluate frontier AI coding models by performing realistic data...  ...distributed systems. Join a sprint-based program where each accepted task is compensated... 

    Obsidian

    San Francisco, CA
    4 days ago
  • $218.5k - $288k

     ...Are We?Postman is the world’s leading API platform, used by more...  ...specializing in Small Language Models and AI Training, you will lead...  ...models.Design, implement, and evaluate model training experiments to...  ...or efficient models.Strong programming skills in Python and familiarity... 
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    2 days ago
  • Anthropic is seeking a Threat Intel Manager to build and lead the Model Exploitation & Fraud team within Threat Intelligence in San Francisco. You will set strategy, hire and guide investigators, and scale systems for fast, high-volume investigations across model distillation... 

    Anthropic

    San Francisco, CA
    4 days ago
  • $85 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco,...  ...frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    a month ago
  •  ...seeking an innovative Quality Engineer for their AI products. This role blends ops, strategy, and...  ...to shape how AI behaves, work with partners in leading labs, and ensure user satisfaction through effective evaluation baselines. Competitive salary and benefits offered... 

    Notion

    San Francisco, CA
    1 day ago
  • $85 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our...  ...Use frontier AI coding agents to complete and evaluate complex engineering tasks. Review model-generated mobile application code for correctness,... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    21 days ago
  • Snap Inc. is seeking a Compensation Program Manager to own end-to-end compensation planning cycles, including...  ..., and roll out programs while championing AI-enabled processes and analytics. You will analyze market trends, lead critical compensation initiatives, and develop... 

    Socket.dev

    San Francisco, CA
    4 days ago
  • MC2 is seeking a Senior Program Design professional to initiate and lead high-impact initiatives within its Strategic Enablement pillars, owning from problem...  ...track effectiveness, and continuously refine based on data and feedback, using AI to #J-18808-Ljbffr Jobtailor

    Jobtailor

    San Francisco, CA
    2 days ago
  • A leading technology firm in San Francisco Bay Area is seeking a GenAI Program Manager to lead strategic initiatives related to AI and machine learning. The ideal candidate will have over 15 years of experience in project management, focusing on integrating AI-driven solutions... 
    Contract work

    Enexus Global Inc.

    San Francisco, CA
    2 days ago
  • Block, Inc. seeks an AI Legal Program Manager to navigate the legal and regulatory landscape of our AI initiatives. You will partner with legal...  ...and services. You will develop governance docs, monitor model deployments, and drive legally compliant AI development practices... 

    Block, Inc.

    San Francisco, CA
    2 days ago
  •  ...experience in digital customer success at a high-growth SaaS company, with proven experience in building successful automated customer programs. Compensation includes a competitive base salary ranging from 180K to 280K, along with generous benefits and a lunch stipend. #J-... 

    Juicebox App, Inc.

    San Francisco, CA
    4 days ago
  • $70 - $75 per hour

    Solomon Page is seeking an organized and proactive AI Enablement & Adoption Coordinator to support the day-to-day operations of enterprise AI learning and adoption programs in San Francisco. You will help drive AI proficiency across the organization by supporting learning... 
    Hourly pay

    Solomon Page

    San Francisco, CA
    3 days ago
  • $290k

    Anthropic is seeking a Technical Program Manager for Cloud Inference in New York. You will drive coordination across multiple teams and manage the launch of AI models on cloud platforms such as Amazon Bedrock and Google Vertex AI. This role requires expertise in technical... 

    Anthropic

    San Francisco, CA
    3 days ago
  • Airbnb, Inc. in San Francisco is seeking a Staff Program Manager, Technical Education to lead engineering education initiatives, including AI education, and to scale technical learning programs across the company. You will design content, manage in-person and virtual events... 

    airbnb, Inc.

    San Francisco, CA
    2 days ago
  •  ...its ML Data Team in San Francisco. This role involves designing evaluation frameworks, managing data operations, and collaborating cross-functionally...  .... Ideal candidates should have over 5 years of experience in AI data operations, proficiency in Python and a strong ability to... 
    Flexible hours

    TwelveLabs

    San Francisco, CA
    3 days ago
  • $127k - $269k

    Figma is seeking a strategic, data-driven Voice of the Customer (VOC) Program Manager to lead a company-wide VOC program. This role is pivotal in surfacing customer insights from various sources and ensuring they are acted upon to improve processes and products. The ideal... 
    Full time
    Remote work

    Figma

    San Francisco, CA
    2 days ago
  • Abridge is seeking an AI Enablement Program Manager in San Francisco to drive AI tool adoption and literacy across the company. This role involves designing effective programs, managing tool lifecycle approval processes, and tracking AI adoption metrics. The ideal candidate... 

    Abridge

    San Francisco, CA
    2 days ago
  • Crossover is hiring a Reading Program Coordinator to lead on-site K-3 literacy workshops across multiple Alpha campuses. You will design and deliver structured-literacy small-group sessions, using AI-adaptive data to drive instruction and track progress weekly. You will... 

    Crossover

    San Francisco, CA
    4 days ago
  • A cutting-edge AI technology firm in San Francisco is seeking an Evaluation Lead to drive the assessment of AI model performance. You will design evaluation methodologies, automate evaluation processes, and oversee various evaluation strategies. The ideal candidate has... 

    SupportFinity

    San Francisco, CA
    1 day ago
  • Charta Health, based in San Francisco, seeks a Technical Program Manager to own the operating system for our engineering organization. You...  ...cross-functional roadmaps, and drive the launch of new AI capabilities across the platform. You will support multiple concurrent... 

    Charta Health

    San Francisco, CA
    3 days ago
  •  ...communities around gaming, entertainment, music, and more. As a Program Manager on the PEP team, you’ll drive employee engagement initiatives...  ..., and Twitch Clubs, while modernizing delivery through AI and delivering impact with measurable outcomes. #J-18808-Ljbffr... 

    Twitch

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Model Evaluation Program Lead. Be the first to apply!