Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Practicing Physician for AI Model Evaluation

$70 - $110 per hour

SaidGig

Role Overview

Shape how advanced AI systems reason about real clinical work. In this senior clinical medicine role, you will partner with an AI research team to evaluate medical knowledge tasks, define high-quality clinical standards, and develop benchmarks that measure meaningful model improvement.

Key Responsibilities

  • Review clinical knowledge tasks and model outputs for missing behaviors, weak reasoning, unsafe recommendations, guideline misalignment, and answers that would not withstand clinical scrutiny.
  • Write detailed instruction specifications and gold-standard solutions for clinical problems.
  • Create new clinical tasks that reflect real-world medical practice.
  • Design challenging clinical evaluations and benchmarks, and help develop medicine-specific skills and tools with the research team.
  • Collaborate with researchers and adjacent-domain specialists to calibrate standards and translate clinical judgment into clear, teachable criteria.

Qualifications

  • MD or DO from an accredited medical school, with completed residency training in a recognized specialty.
  • At least 4 years of post-residency clinical practice. Residency and fellowship training do not count toward this requirement.
  • Active, unrestricted license to practice medicine in at least one U.S. state and board certification in your specialty.
  • Established expertise in a clinical specialty, such as internal medicine, oncology, radiology, emergency medicine, surgery, psychiatry, or a medical subspecialty.
  • Senior-level clinical leadership or decision ownership, such as Attending Physician, Medical Director, Division Chief, Associate or Full Professor, or Chief Medical Officer.
  • Professional, hands-on experience using large language models and the judgment to distinguish sound clinical reasoning from plausible but incorrect answers.
  • Excellent written communication and the ability to provide precise, well-structured feedback.
  • Experience in utilization management, clinical informatics, or medical affairs is a plus.

Work Terms

  • Full-time W-2 employment, with a reliable commitment of 40 hours per week for an initial 6-month engagement.
  • Hybrid role based in the Bay Area, California. You must live in the Bay Area and be available to work on-site with the client team multiple days per week when required.
  • This is not a remote role. Candidates relocating to the Bay Area must do so at their own expense; relocation assistance is not available.
  • You will work within the client’s tools alongside its research team and receive client-issued accounts and equipment.

Compensation

$70 to $110 per hour.

Eligibility

Qualified applicants are considered without discrimination based on legally protected characteristics. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans during the application process.

Application Process

Applicants may discover this opportunity through an online talent platform. Employment, onboarding, payroll, benefits, and compliance are managed by the employer of record, while the successful candidate is placed directly within the client team.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Practicing Physician for AI Model Evaluation in California vacancy
  • $136.44k - $265.11k

    We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world waits...  ...structured data, and run AI agents and models directly in their workflows. Over 200,000...  ...our work here.You’ll build the datasets, evaluations, and systems that help close that gap.... 
    Suggested
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    2 days ago
  • $100 - $150 per hour

     ...senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You...  ...high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal specialist with deep subject-matter... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Suggested
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab'...  ...Overview This senior, specialized role turns practical scientific expertise into high-quality... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $70 per hour

     ...interactions) Develop temporal graph modeling techniques to capture time-varying multi-...  ...Demonstrated research background or practical experience in Knowledge Graphs, Graph Algorithms...  ...of published papers in top-tier AI/ML, data mining, or computer vision conferences... 
    Suggested
    Hourly pay
    Full time
    Internship
    Summer internship
    Relocation package

    Waymo

    Mountain View, CA
    11 hours ago
  • $400 per month

    About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering... 

    Mercor Inc

    San Francisco, CA
    5 days ago
  • $100 - $150 per hour

     ...authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role...  ...or more years of substantive post-qualification legal practice at a reputable institution, for example an established... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Local area
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $70 - $110 per hour

     ...Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work...  ...quality materials reasoning looks like in practice and ensure model outputs can withstand technical... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $75 - $115 per hour

     ...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific...  ...drug development reasoning looks like in practice. Key Responsibilities Review... 
    Hourly pay
    Full time
    Contract work
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $65 - $105 per hour

     ...Overview Help advance frontier AI systems by applying deep...  ...benchmarks used to assess how models reason about real software and...  ...program management teams, turning practical expertise into clear...  ...Design rigorous engineering evaluation sets and contribute to engineering... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $65 - $105 per hour

     ...Help improve how frontier AI models reason about real-world life sciences research. In this...  ...will apply deep scientific judgment to evaluate research tasks and model outputs, define...  ...define tasks that reflect real research practice. Design challenging domain-specific evaluation... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $190k - $250k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...developing large-scale generative world models that learn to predict realistic, physically...  ...between camera, LiDAR, and radar outputsDesign evaluation frameworks that measure world model... 
    Temporary work
    Work at office
    Visa sponsorship

    Kodiak Robotics

    Mountain View, CA
    5 days ago
  • $100 - $150 per hour

     ...Role Overview Help shape how next-generation AI models perform real financial work by providing deep, practical finance expertise to a GenAI research team. You will...  ...depth: Design challenging finance tasks and evaluation sets, and collaborate with researchers to build... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  •  ...developed the world's first real-time speech AI platform capable of accent translation,...  ...ARR. Our team combines deep expertise in model innovation and systems engineering with a...  ...all of Sanas's model families, build the evaluation infrastructure to measure it rigorously,... 

    Sanas.AI Inc.

    Palo Alto, CA
    4 days ago
  • $60 - $90 per hour

     ...character animation expertise to shape how a performance transfer model is evaluated, ensuring actor timing, emotional intent, and subtle physical...  ...pool that scores model outputs for a leading generative AI research effort. Key Responsibilities Define evaluation... 
    Hourly pay
    Part time
    Freelance

    SaidGig

    Playa Vista, CA
    1 day ago
  • $174.72k - $295.68k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  ...with strong expertise in generative modeling and large-scale deep learning systems, along...  ...role, you will research, implement, and evaluate world models that learn the dynamics of... 
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...the next generation of Physical AI, spanning data generation, large-scale multimodal model training, robotics simulation...  ...or other post-training methods, evaluation, and model optimization.Strong...  ...including in our hiring and promotion practices) on the basis of race, religion... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $219k - $351k

     ...becoming a memory-bandwidth business. As models scale past what any single GPU can hold —...  ...that treat memory as the core product of AI inference, not an afterthought.We are looking...  ...managers to ensure every candidate is evaluated fairly and holistically.Recruiting Agency... 
    Work at office
    Remote work
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    3 days ago
  • $272k - $431.25k

     ...we’re generating it! Our world model team is pushing the boundaries of multimodal AI, robotics, and world foundation...  ...Research Manager to lead world-model evaluation and benchmarking across NVIDIA’s...  ...in our hiring and promotion practices) on the basis of race, religion,... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $174.72k - $295.68k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  .../ Research Scientist to drive the modeling and algorithmic development of XPENG’s next...  ...DDP, etc.).Conduct systematic ablation, evaluation, and visualization of model behavior across... 
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago
  •  ...providing independent assurance and evaluating the company's risk management...  ...for data scientists and AI developers who will power our...  ...in frameworks for auditing models, including criteria like robustness...  ...knowledge in data analytics practices, machine learning, AI, and... 

    TikTok

    Los Angeles, CA
    3 days ago
  • $220k - $320k

     ...Inference.net]( trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match...  ...everything end-to-end: distillation, training, evaluation, and planet-scale hosting. We are a well-funded ten... 
    Full time
    Work at office

    Inference Corp

    San Francisco, CA
    1 day ago
  • $172.43k - $230.95k

     ...intelligence . As the only vertically integrated AI infrastructure company built from the...  ...The Senior Software Engineer for the AI Model Lifecycle team will play a crucial role...  ...management: versioning, lineage, evaluation, and reproducible fine-tuning at scale.... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $220k - $300k

     ...About OpenAI Foundation AI is already changing how people work, learn, and access care...  ...of the work of our mission. Advanced AI models will also present new challenges that are...  ...improving how models behave (testing and evaluating them independently, red-teaming them,... 
    Full time

    OpenAI Foundation

    San Francisco, CA
    3 days ago
  •  ...focus on developing and applying AI and machine learning methods...  ...with an emphasis on rigorous evaluation in real-world settings. The...  ...clinical utility of ML/AI models intended for translational impact...  ..., platforms, and engineering practices used across data and software... 

    Stanford University

    Stanford, CA
    2 days ago
  • $184k - $287.5k

     ...is redefining what is possible with AI, and the Relational Foundation Model team is helping lead that...  ...models: you will design, build, and evaluate novel Transformer and graph neural...  ...including in our hiring and promotion practices) on the basis of race, religion, color... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...into the unlimited potential of AI to define the next era of...  ...opportunity to build a groundbreaking model customization and deployment...  ...through fine-tuning, evaluation, deployment, and compliance flawlessly...  ...in our hiring and promotion practices) on the basis of race, religion... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $300k - $333k

     ...tone, and behavior with the model specifications and taxonomy,...  ...scaling processes for consistent evaluations.Dive deep into technical...  ...technical field or equivalent practical experience.10 years of experience...  ...for ambiguity inherent in AI research and operates at a fast... 

    Google

    Mountain View, CA
    1 day ago
  • $110 - $150 per hour

     ...Role Overview Help advance frontier AI models by bringing professional finance judgment to the evaluation, design, and improvement of financial knowledge-work tasks...  ...financial reasoning looks like in real-world practice. Key Responsibilities Review finance tasks... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $218.5k - $288k

     ...Scientist specializing in Small Language Models and AI Training, you will lead research and...  ...performance language models tailored for practical applications. You will work closely...  ...language models.Design, implement, and evaluate model training experiments to improve performance... 
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Practicing Physician for AI Model Evaluation. Be the first to apply!