Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Practicing Physician for AI Model Evaluation

$70 - $110 per hour

SaidGig

Role Overview

Shape how advanced AI systems reason about real clinical work. In this senior clinical medicine role, you will partner with an AI research team to evaluate medical knowledge tasks, define high-quality clinical standards, and develop benchmarks that measure meaningful model improvement.

Key Responsibilities

  • Review clinical knowledge tasks and model outputs for missing behaviors, weak reasoning, unsafe recommendations, guideline misalignment, and answers that would not withstand clinical scrutiny.
  • Write detailed instruction specifications and gold-standard solutions for clinical problems.
  • Create new clinical tasks that reflect real-world medical practice.
  • Design challenging clinical evaluations and benchmarks, and help develop medicine-specific skills and tools with the research team.
  • Collaborate with researchers and adjacent-domain specialists to calibrate standards and translate clinical judgment into clear, teachable criteria.

Qualifications

  • MD or DO from an accredited medical school, with completed residency training in a recognized specialty.
  • At least 4 years of post-residency clinical practice. Residency and fellowship training do not count toward this requirement.
  • Active, unrestricted license to practice medicine in at least one U.S. state and board certification in your specialty.
  • Established expertise in a clinical specialty, such as internal medicine, oncology, radiology, emergency medicine, surgery, psychiatry, or a medical subspecialty.
  • Senior-level clinical leadership or decision ownership, such as Attending Physician, Medical Director, Division Chief, Associate or Full Professor, or Chief Medical Officer.
  • Professional, hands-on experience using large language models and the judgment to distinguish sound clinical reasoning from plausible but incorrect answers.
  • Excellent written communication and the ability to provide precise, well-structured feedback.
  • Experience in utilization management, clinical informatics, or medical affairs is a plus.

Work Terms

  • Full-time W-2 employment, with a reliable commitment of 40 hours per week for an initial 6-month engagement.
  • Hybrid role based in the Bay Area, California. You must live in the Bay Area and be available to work on-site with the client team multiple days per week when required.
  • This is not a remote role. Candidates relocating to the Bay Area must do so at their own expense; relocation assistance is not available.
  • You will work within the client’s tools alongside its research team and receive client-issued accounts and equipment.

Compensation

$70 to $110 per hour.

Eligibility

Qualified applicants are considered without discrimination based on legally protected characteristics. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans during the application process.

Application Process

Applicants may discover this opportunity through an online talent platform. Employment, onboarding, payroll, benefits, and compliance are managed by the employer of record, while the successful candidate is placed directly within the client team.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Practicing Physician for AI Model Evaluation in California vacancy
  • $136.44k - $265.11k

    We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world waits...  ...structured data, and run AI agents and models directly in their workflows. Over 200,000...  ...our work here.You’ll build the datasets, evaluations, and systems that help close that gap.... 
    Suggested
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    7 days ago
  • $65 - $105 per hour

     ...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world software development...  ...team to define high-quality engineering work, evaluate model performance, and turn expert practice into clear standards and benchmarks. Key... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $60 - $90 per hour

     ...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,...  ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities... 
    Suggested
    Hourly pay
    Part time

    SaidGig

    Playa Vista, CA
    a month ago
  • $100 - $150 per hour

     ...authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role...  ...or more years of substantive post-qualification legal practice at a reputable institution, for example an established... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Local area
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $75 - $115 per hour

     ...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific...  ...drug development reasoning looks like in practice. Key Responsibilities Review... 
    Suggested
    Hourly pay
    Full time
    Contract work
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    2 days ago
  • $70 - $110 per hour

     ...Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work...  ...quality materials reasoning looks like in practice and ensure model outputs can withstand technical... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    1 day ago
  • $65 - $105 per hour

     ...Help improve how frontier AI models reason about real-world life sciences research. In this...  ...will apply deep scientific judgment to evaluate research tasks and model outputs, define...  ...define tasks that reflect real research practice. Design challenging domain-specific evaluation... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    2 days ago
  • $100 - $150 per hour

     ...Role Overview Help shape how next-generation AI models perform real financial work by providing deep, practical finance expertise to a GenAI research team. You will...  ...depth: Design challenging finance tasks and evaluation sets, and collaborate with researchers to build... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    2 days ago
  • $60 - $90 per hour

     ...character animation expertise to shape how a performance transfer model is evaluated, ensuring actor timing, emotional intent, and subtle physical...  ...pool that scores model outputs for a leading generative AI research effort. Key Responsibilities Define evaluation... 
    Hourly pay
    Part time
    Freelance

    SaidGig

    Playa Vista, CA
    1 day ago
  • $218.5k - $288k

     ...Scientist specializing in Small Language Models and AI Training, you will lead research and...  ...performance language models tailored for practical applications. You will work closely...  ...language models.Design, implement, and evaluate model training experiments to improve performance... 
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    a month ago
  •  ...Description Job Description As the Manager of Model Validation & Verification (VnV) for...  ...and data science team responsible for evaluating, benchmarking, and validating the machine...  ...experience in robotics, autonomous systems, or AI/ML. Domain Knowledge: Strong... 
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    5 days ago
  • $168k - $268.4k

     ...invite you to join us. Where AI Meets Medicine: Build the Future...  ...foundation and frontier AI models trained on Lilly data at scale...  ..., you will design, train, and evaluate foundation models that advance...  ...balancing scientific innovation with practical application.What You Should... 
    Remote work
    Flexible hours
    2 days per week

    Eli Lilly and Company

    San Francisco, CA
    11 days ago
  • $190k - $250k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...developing large-scale generative world models that learn to predict realistic, physically...  ...camera, LiDAR, and radar outputs Design evaluation frameworks that measure world model... 
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    18 days ago
  • $220k - $320k

     ...Inference.net]( trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match...  ...everything end-to-end: distillation, training, evaluation, and planet-scale hosting. We are a well-funded ten... 
    Work at office

    Inference

    San Francisco, CA
    1 day ago
  •  ...TikTok is seeking data scientists and AI developers to power continuous auditing and risk identification across verticals. You...  ...to build analytics products for the audit team. Join a team evaluating model lifecycles, biases, and security risks, while delivering scalable... 

    Jobleads-US

    San Jose, CA
    2 days ago
  • $272k - $431.25k

     ...we’re generating it! Our world model team is pushing the boundaries of multimodal AI, robotics, and world foundation...  ...Research Manager to lead world-model evaluation and benchmarking across NVIDIA’s...  ...in our hiring and promotion practices) on the basis of race, religion,... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $136.8k - $277.2k

     ...providing independent assurance and evaluating the company\'s risk management...  ...for data scientists and AI developers who will power our...  ...team. Responsibilities Model Evaluation & Audit Frameworks:...  ...expand knowledge in data analytics practices, machine learning, AI, and... 
    Temporary work
    Local area

    Jobleads-US

    San Jose, CA
    2 days ago
  • $60 - $100 per hour

     ...Help advance frontier AI systems by bringing practicing insurance and actuarial judgment into the evaluation, training, and improvement of models used for real insurance work. You will work closely with an AI research team to define high-quality insurance reasoning, identify... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  •  ...the Institute of Foundation Models  We are a dedicated research...  ...nurture the next generation of AI builders, and drive...  ...pipelines, experimentation, and evaluation workflows.  ~ This role balances...  ...and drive best practices in reliability, scalability,... 
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    24 days ago
  • $219k - $351k

     ...becoming a memory-bandwidth business. As models scale past what any single GPU can hold —...  ...that treat memory as the core product of AI inference, not an afterthought.\nWe are looking...  ...managers to ensure every candidate is evaluated fairly and holistically.\n Recruiting... 
    Work at office
    Remote work
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    5 days ago
  • $152k - $241.5k

     ...the next generation of Physical AI, spanning data generation, large-scale multimodal model training, robotics simulation...  ...or other post-training methods, evaluation, and model optimization.Strong...  ...including in our hiring and promotion practices) on the basis of race, religion... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $174.72k - $295.68k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  ...with strong expertise in generative modeling and large-scale deep learning systems, along...  ...role, you will research, implement, and evaluate world models that learn the dynamics of... 
    Full time

    XPENG Motors

    Santa Clara, CA
    24 days ago
  • $300k - $333k

     ...tone, and behavior with the model specifications and taxonomy,...  ...scaling processes for consistent evaluations.Dive deep into technical...  ...technical field or equivalent practical experience.10 years of experience...  ...for ambiguity inherent in AI research and operates at a fast... 

    Google

    Mountain View, CA
    7 days ago
  • $215.28k - $364.32k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  .../ Research Scientist to drive the modeling and algorithmic development of XPENG’s next...  ..., etc.). Conduct systematic ablation, evaluation, and visualization of model behavior across... 
    Full time

    XPENG

    California
    22 days ago
  • $184k - $287.5k

     ...is redefining what is possible with AI, and the Relational Foundation Model team is helping lead that...  ...models: you will design, build, and evaluate novel Transformer and graph neural...  ...including in our hiring and promotion practices) on the basis of race, religion, color... 
    Full time

    Nvidia

    Santa Clara, CA
    19 days ago
  • $184k - $287.5k

     ...into the unlimited potential of AI to define the next era of...  ...opportunity to build a groundbreaking model customization and deployment...  ...through fine-tuning, evaluation, deployment, and compliance flawlessly...  ...in our hiring and promotion practices) on the basis of race, religion... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  •  ...focus on developing and applying AI and machine learning methods...  ...with an emphasis on rigorous evaluation in real-world settings. The...  ...clinical utility of ML/AI models intended for translational impact...  ..., platforms, and engineering practices used across data and software... 

    Stanford University

    Stanford, CA
    13 days ago
  • $150k

     ...Description SpaceXAI's mission is to create AI systems that can accurately understand...  ...THE ROLE: You will join the Grok Voice Model team to help build the world's best voice...  ...to enable high-quality model training and evaluation. Work on pre-training and post-... 
    Temporary work

    SpaceXAI

    Palo Alto, CA
    18 days ago
  •  ...healthcare and scientific discovery to powering AI and the technologies people rely on every...  ...: We are seeking highly motivated AI Model Optimization & Software Engineer Interns/...  ...Runtime, vLLM, or SGLang. Develop or evaluate parallel and distributed computing methods... 
    Full time
    Summer work
    Internship
    Summer internship
    Worldwide

    AMD

    San Jose, CA
    2 days ago
  • $237.6k - $318.24k

     ...intelligence . As the only vertically integrated AI infrastructure company built from the...  ...Staff Software Engineer for the AI Model Lifecycle team will play a crucial role in...  ...experiment management: versioning, lineage, evaluation, and reproducible fine-tuning at scale.... 
    Temporary work

    Crusoe

    San Francisco, CA
    26 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Practicing Physician for AI Model Evaluation. Be the first to apply!