Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Practicing Clinician for AI Model Evaluation

$70 - $110 per hour

SaidGig

Role Overview

Help a leading AI research team improve how advanced AI models reason about real clinical work. In this hybrid, full-time role, you will apply practicing-clinician judgment to define high-quality clinical tasks, model answers, and evaluation standards alongside research and program management teams.

Key Responsibilities

  • Review clinical knowledge tasks and model outputs for missing behaviors, weak reasoning, unsafe recommendations, departures from clinical guidelines, and plausible-sounding answers that do not meet clinical standards.
  • Write detailed instruction specifications and gold-standard solutions for clinical problems.
  • Create clinical tasks that reflect real-world medical practice.
  • Design challenging benchmarks and evaluation sets to measure model improvement.
  • Partner with researchers to develop medicine-specific capabilities and tools.
  • Calibrate standards with researchers and adjacent-domain specialists, turning implicit clinical judgment into clear, teachable criteria.

Qualifications

  • MD or DO from an accredited medical school, with completed residency training in a recognized specialty.
  • At least 4 years of post-residency clinical practice. Residency and fellowship training do not count toward this requirement.
  • Active, unrestricted medical license in at least one U.S. state and board certification in your specialty.
  • Established expertise in a clinical specialty, such as internal medicine, oncology, radiology, emergency medicine, surgery, psychiatry, or a medical subspecialty.
  • Senior clinical progression, such as Attending Physician, Medical Director, Division Chief, Associate or full Professor, or Chief Medical Officer, with meaningful ownership of clinical decisions.
  • Hands-on professional use of large language models and the ability to distinguish sound clinical reasoning from convincing but incorrect answers.
  • Excellent written communication and the ability to provide precise, structured feedback.
  • Experience in utilization management, clinical informatics, or medical affairs is a plus.

Work Terms

  • Full-time W-2 hourly employment, with an initial 6-month commitment and reliable availability for 40 hours per week.
  • Hybrid role based in the Bay Area, California. You must live in the Bay Area and be available to work on-site with the client team multiple days per week when required.
  • This is not a remote position. Candidates outside the Bay Area must relocate at their own expense before the engagement begins; relocation assistance is not provided.
  • You will work within the client’s tools alongside internal research teams and receive client-issued accounts and equipment.
  • The position is a structured, role-based placement within an enterprise team, not a freelance engagement.

Compensation

$70 to $110 per hour.

Application Process

Opportunities may be discovered through an online job platform. Employment, onboarding, payroll, benefits, and compliance are managed by the employer of record for the engagement.

Equal Opportunity

Equal employment opportunity is provided without discrimination based on any legally protected characteristic. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans throughout the application process.

Vacancy posted 20 hours ago
Similar jobs that could be interesting for youBased on the Practicing Clinician for AI Model Evaluation in California vacancy
  • $224k - $356.5k

     ...tapping into the unlimited potential of AI to define the next era of computing. An...  ...Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful...  ..., shaping the roadmap, and sharing best practices.Work alongside model training, inference... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $70 - $110 per hour

     ...Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work...  ...quality materials reasoning looks like in practice and ensure model outputs can withstand technical... 
    Suggested
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    20 hours ago
  • $300k - $320k

     ...role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be...  ...partners and industry standards bodies to align our evaluation practices with emerging best practices in responsible AI development... 
    Suggested
    Work at office
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    San Francisco, CA
    4 days ago
  • $60 - $100 per hour

     ...Role Overview Help advance frontier AI models by bringing senior insurance and actuarial judgment to the evaluation of real-world insurance work. You will work directly...  ...Create tasks that reflect real insurance practice, along with challenging benchmarks and evaluation... 
    Suggested
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    20 hours ago
  • $184.7k - $324.8k

    Research Scientist / Engineer, Foundation Model Evaluation Cupertino, California, United States...  ...Qualifications 3+ years of experience in AI model evaluation, NLP, or a related area...  ...ability to translate research insights into practical implementations Strong experimental... 
    Suggested
    Relocation

    Apple Inc.

    Cupertino, CA
    1 day ago
  • $238k - $302k

     ...in simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language Models...  ...with software design principles, coding best practices, testing methodologies, and version control software... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Weekday work

    Mercor Inc

    San Francisco, CA
    2 days ago
  • Apple Inc. is seeking a Research Scientist/Engineer to design evaluation systems for foundation models powering Apple products. You will work hands‑on across evaluation design, experimentation, and cross‑team collaboration to drive model improvement and product quality... 

    Apple Inc.

    Cupertino, CA
    1 day ago
  • $50 - $75 per hour

    A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in... 
    Hourly pay
    Contract work

    Mercor

    San Francisco, CA
    1 day ago
  • Job Description - Member of Technical Staff (Language Model Evaluations) Location: San Francisco (preferred), Sydney, Melbourne, Brisbane About...  ...Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises... 

    Artificial Analysis, Inc.

    San Francisco, CA
    2 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    5 days ago
  • Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM... 
    Weekday work

    Mercor

    San Francisco, CA
    3 days ago
  • $350k

     ...Our first goal is to democratize frontier AI R&D across scientific disciplines. We...  ...AI research company and training our own models end-to-end. Our work spans areas such as...  ...looking for a research engineer to build the evaluation infrastructure that tells us whether our... 

    Mirendil

    San Francisco, CA
    1 day ago
  • A leading AI company is seeking a legal professional for a contractor role focused on evaluating AI model outputs in legal contexts. Candidates must hold a Juris Doctor (J.D.) and have more than 3 years of experience in law. The role involves reviewing complex legal hypotheticals... 
    For contractors
    10 hours per week

    Turing

    Los Angeles, CA
    2 days ago
  • $400 per month

    About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows... 

    Obsidian

    San Francisco, CA
    3 days ago
  • $218.5k - $288k

     ...Scientist specializing in Small Language Models and AI Training, you will lead research and...  ...performance language models tailored for practical applications. You will work closely...  ...language models.Design, implement, and evaluate model training experiments to improve performance... 
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    4 days ago
  • $65 - $105 per hour

     ...Help improve how frontier AI models reason about real-world life sciences research. In this...  ...will apply deep scientific judgment to evaluate research tasks and model outputs, define...  ...define tasks that reflect real research practice. Design challenging domain-specific evaluation... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $100 - $150 per hour

     ...Role Overview Help shape how next-generation AI models perform real financial work by providing deep, practical finance expertise to a GenAI research team. You will...  ...depth: Design challenging finance tasks and evaluation sets, and collaborate with researchers to build... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    2 days ago
  • $195.2k - $262.2k

     ...infrastructure for the global AI economy. We are building a full...  ...and enterprises from data and model training through to production...  ..., task environments, and evaluation sets for reasoning, coding, tool...  ..., and failure analysis. Practical understanding of modern LLM behavior... 
    Full time
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    3 days ago
  • $272k - $431.25k

     ...we’re generating it! Our world model team is pushing the boundaries of multimodal AI, robotics, and world foundation...  ...Research Manager to lead world-model evaluation and benchmarking across NVIDIA’s...  ...in our hiring and promotion practices) on the basis of race, religion,... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...providing independent assurance and evaluating the company's risk management...  ...for data scientists and AI developers who will power our...  ...in frameworks for auditing models, including criteria like robustness...  ...knowledge in data analytics practices, machine learning, AI, and... 

    TikTok

    Los Angeles, CA
    1 day ago
  • $95.68k - $164.32k

     ...transformer architectures, foundation models, and telemetry-driven AI to impact millions of ArcGIS users...  ...infrastructureDesign and implement evaluation frameworks that measure model quality...  ...documents, experiment reports, and best-practice guidance for model development and... 
    Worldwide

    ESRI

    Redlands, CA
    2 days ago
  • $400 per month

     ...About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering... 

    Mercor Inc

    San Francisco, CA
    1 day ago
  •  ...San Francisco is seeking an innovative Quality Engineer for their AI products. This role blends ops, strategy, and analytics to...  ...in leading labs, and ensure user satisfaction through effective evaluation baselines. Competitive salary and benefits offered, with a focus... 

    Notion

    San Francisco, CA
    2 days ago
  • $85 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our...  ...Use frontier AI coding agents to complete and evaluate complex engineering tasks. Review model-generated mobile application code for correctness, quality... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    23 days ago
  • $184k - $287.5k

     ...is redefining what is possible with AI, and the Relational Foundation Model team is helping lead that...  ...models: you will design, build, and evaluate novel Transformer and graph neural...  ...including in our hiring and promotion practices) on the basis of race, religion, color... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $175k - $215k

     ...states. The mission of the Waymo AI Foundations team is to develop...  ...demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you...  ...field of study, or equivalent practical experience Proficiency in... 
    Full time
    Remote work

    Waymo

    Mountain View, CA
    4 days ago
  • $192k - $278k

    Lead model releases for Search, evaluating DeepMind release applicants against strict quality bars to determine...  ...as Search evolves into a fully AI-enabled product.Design and execute end...  ...in a technical field, or equivalent practical experience. 8 years of experience in... 
    Shift work

    Google

    Mountain View, CA
    2 days ago
  • $75 - $115 per hour

     ...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific...  ...drug development reasoning looks like in practice. Key Responsibilities Review... 
    Hourly pay
    Full time
    Contract work
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    20 hours ago
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes, and... 

    Mercor Inc

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Practicing Clinician for AI Model Evaluation. Be the first to apply!