Practicing Physician for AI Model Evaluation
$70 - $110 per hourSaidGig
Role Overview
Shape how advanced AI systems reason about real clinical work. In this senior clinical medicine role, you will partner with an AI research team to evaluate medical knowledge tasks, define high-quality clinical standards, and develop benchmarks that measure meaningful model improvement.
Key Responsibilities
- Review clinical knowledge tasks and model outputs for missing behaviors, weak reasoning, unsafe recommendations, guideline misalignment, and answers that would not withstand clinical scrutiny.
- Write detailed instruction specifications and gold-standard solutions for clinical problems.
- Create new clinical tasks that reflect real-world medical practice.
- Design challenging clinical evaluations and benchmarks, and help develop medicine-specific skills and tools with the research team.
- Collaborate with researchers and adjacent-domain specialists to calibrate standards and translate clinical judgment into clear, teachable criteria.
Qualifications
- MD or DO from an accredited medical school, with completed residency training in a recognized specialty.
- At least 4 years of post-residency clinical practice. Residency and fellowship training do not count toward this requirement.
- Active, unrestricted license to practice medicine in at least one U.S. state and board certification in your specialty.
- Established expertise in a clinical specialty, such as internal medicine, oncology, radiology, emergency medicine, surgery, psychiatry, or a medical subspecialty.
- Senior-level clinical leadership or decision ownership, such as Attending Physician, Medical Director, Division Chief, Associate or Full Professor, or Chief Medical Officer.
- Professional, hands-on experience using large language models and the judgment to distinguish sound clinical reasoning from plausible but incorrect answers.
- Excellent written communication and the ability to provide precise, well-structured feedback.
- Experience in utilization management, clinical informatics, or medical affairs is a plus.
Work Terms
- Full-time W-2 employment, with a reliable commitment of 40 hours per week for an initial 6-month engagement.
- Hybrid role based in the Bay Area, California. You must live in the Bay Area and be available to work on-site with the client team multiple days per week when required.
- This is not a remote role. Candidates relocating to the Bay Area must do so at their own expense; relocation assistance is not available.
- You will work within the client’s tools alongside its research team and receive client-issued accounts and equipment.
Compensation
$70 to $110 per hour.
Eligibility
Qualified applicants are considered without discrimination based on legally protected characteristics. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans during the application process.
Application Process
Applicants may discover this opportunity through an online talent platform. Employment, onboarding, payroll, benefits, and compliance are managed by the employer of record, while the successful candidate is placed directly within the client team.
$136.44k - $265.11k
We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world waits... ...structured data, and run AI agents and models directly in their workflows. Over 200,000... ...our work here.You’ll build the datasets, evaluations, and systems that help close that gap....SuggestedWork at officeLocal areaMonday to FridayShift work$65 - $105 per hour
...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world software development... ...team to define high-quality engineering work, evaluate model performance, and turn expert practice into clear standards and benchmarks. Key...SuggestedHourly payFull timeFreelanceInternshipLive inRelocationRelocation package$60 - $90 per hour
...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,... ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities...SuggestedHourly payPart time$100 - $150 per hour
...authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role... ...or more years of substantive post-qualification legal practice at a reputable institution, for example an established...SuggestedHourly payFull timeFreelanceInternshipLive inLocal areaRelocationRelocation package$75 - $115 per hour
...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific... ...drug development reasoning looks like in practice. Key Responsibilities Review...SuggestedHourly payFull timeContract workLive inRelocationRelocation package$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work... ...quality materials reasoning looks like in practice and ensure model outputs can withstand technical...Hourly payFull timeLive inRelocationRelocation package$65 - $105 per hour
...Help improve how frontier AI models reason about real-world life sciences research. In this... ...will apply deep scientific judgment to evaluate research tasks and model outputs, define... ...define tasks that reflect real research practice. Design challenging domain-specific evaluation...Hourly payFull timeLive inRelocationRelocation package$100 - $150 per hour
...Role Overview Help shape how next-generation AI models perform real financial work by providing deep, practical finance expertise to a GenAI research team. You will... ...depth: Design challenging finance tasks and evaluation sets, and collaborate with researchers to build...Hourly payFull timeLive inRelocationRelocation package$60 - $90 per hour
...character animation expertise to shape how a performance transfer model is evaluated, ensuring actor timing, emotional intent, and subtle physical... ...pool that scores model outputs for a leading generative AI research effort. Key Responsibilities Define evaluation...Hourly payPart timeFreelance$218.5k - $288k
...Scientist specializing in Small Language Models and AI Training, you will lead research and... ...performance language models tailored for practical applications. You will work closely... ...language models.Design, implement, and evaluate model training experiments to improve performance...Work at officeFlexible hours3 days per week- ...Description Job Description As the Manager of Model Validation & Verification (VnV) for... ...and data science team responsible for evaluating, benchmarking, and validating the machine... ...experience in robotics, autonomous systems, or AI/ML. Domain Knowledge: Strong...Temporary workRelocation package
$168k - $268.4k
...invite you to join us. Where AI Meets Medicine: Build the Future... ...foundation and frontier AI models trained on Lilly data at scale... ..., you will design, train, and evaluate foundation models that advance... ...balancing scientific innovation with practical application.What You Should...Remote workFlexible hours2 days per week$190k - $250k
...has developed an artificial intelligence (AI) powered technology stack purpose-built... ...developing large-scale generative world models that learn to predict realistic, physically... ...camera, LiDAR, and radar outputs Design evaluation frameworks that measure world model...Temporary workWork at officeVisa sponsorshipFlexible hours$220k - $320k
...Inference.net]( trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The models we train match... ...everything end-to-end: distillation, training, evaluation, and planet-scale hosting. We are a well-funded ten...Work at office- ...TikTok is seeking data scientists and AI developers to power continuous auditing and risk identification across verticals. You... ...to build analytics products for the audit team. Join a team evaluating model lifecycles, biases, and security risks, while delivering scalable...
$272k - $431.25k
...we’re generating it! Our world model team is pushing the boundaries of multimodal AI, robotics, and world foundation... ...Research Manager to lead world-model evaluation and benchmarking across NVIDIA’s... ...in our hiring and promotion practices) on the basis of race, religion,...Full time$136.8k - $277.2k
...providing independent assurance and evaluating the company\'s risk management... ...for data scientists and AI developers who will power our... ...team. Responsibilities Model Evaluation & Audit Frameworks:... ...expand knowledge in data analytics practices, machine learning, AI, and...Temporary workLocal area$60 - $100 per hour
...Help advance frontier AI systems by bringing practicing insurance and actuarial judgment into the evaluation, training, and improvement of models used for real insurance work. You will work closely with an AI research team to define high-quality insurance reasoning, identify...Hourly payFull timeLive inRelocationRelocation package- ...the Institute of Foundation Models We are a dedicated research... ...nurture the next generation of AI builders, and drive... ...pipelines, experimentation, and evaluation workflows. ~ This role balances... ...and drive best practices in reliability, scalability,...Visa sponsorship
$219k - $351k
...becoming a memory-bandwidth business. As models scale past what any single GPU can hold —... ...that treat memory as the core product of AI inference, not an afterthought.\nWe are looking... ...managers to ensure every candidate is evaluated fairly and holistically.\n Recruiting...Work at officeRemote workFlexible hours$152k - $241.5k
...the next generation of Physical AI, spanning data generation, large-scale multimodal model training, robotics simulation... ...or other post-training methods, evaluation, and model optimization.Strong... ...including in our hiring and promotion practices) on the basis of race, religion...Full time$174.72k - $295.68k
...forefront of innovation, integrating advanced AI and autonomous driving technologies into... ...with strong expertise in generative modeling and large-scale deep learning systems, along... ...role, you will research, implement, and evaluate world models that learn the dynamics of...Full time$300k - $333k
...tone, and behavior with the model specifications and taxonomy,... ...scaling processes for consistent evaluations.Dive deep into technical... ...technical field or equivalent practical experience.10 years of experience... ...for ambiguity inherent in AI research and operates at a fast...$215.28k - $364.32k
...forefront of innovation, integrating advanced AI and autonomous driving technologies into... .../ Research Scientist to drive the modeling and algorithmic development of XPENG’s next... ..., etc.). Conduct systematic ablation, evaluation, and visualization of model behavior across...Full time$184k - $287.5k
...is redefining what is possible with AI, and the Relational Foundation Model team is helping lead that... ...models: you will design, build, and evaluate novel Transformer and graph neural... ...including in our hiring and promotion practices) on the basis of race, religion, color...Full time$184k - $287.5k
...into the unlimited potential of AI to define the next era of... ...opportunity to build a groundbreaking model customization and deployment... ...through fine-tuning, evaluation, deployment, and compliance flawlessly... ...in our hiring and promotion practices) on the basis of race, religion...Full time- ...focus on developing and applying AI and machine learning methods... ...with an emphasis on rigorous evaluation in real-world settings. The... ...clinical utility of ML/AI models intended for translational impact... ..., platforms, and engineering practices used across data and software...
$150k
...Description SpaceXAI's mission is to create AI systems that can accurately understand... ...THE ROLE: You will join the Grok Voice Model team to help build the world's best voice... ...to enable high-quality model training and evaluation. Work on pre-training and post-...Temporary work- ...healthcare and scientific discovery to powering AI and the technologies people rely on every... ...: We are seeking highly motivated AI Model Optimization & Software Engineer Interns/... ...Runtime, vLLM, or SGLang. Develop or evaluate parallel and distributed computing methods...Full timeSummer workInternshipSummer internshipWorldwide
$237.6k - $318.24k
...intelligence . As the only vertically integrated AI infrastructure company built from the... ...Staff Software Engineer for the AI Model Lifecycle team will play a crucial role in... ...experiment management: versioning, lineage, evaluation, and reproducible fine-tuning at scale....Temporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Practicing Physician for AI Model Evaluation. Be the first to apply!
- general practitioner California
- general physician California
- family physician California
- general surgery physician assistant California
- new graduate family nurse practitioner California
- physician family medicine California
- outpatient family medicine physician California
- family medicine physician assistant California
- physician family practice California
- family nurse practitioner California




