Practicing Clinician for AI Model Evaluation
$70 - $110 per hourSaidGig
Role Overview
Help a leading AI research team improve how advanced AI models reason about real clinical work. In this hybrid, full-time role, you will apply practicing-clinician judgment to define high-quality clinical tasks, model answers, and evaluation standards alongside research and program management teams.
Key Responsibilities
- Review clinical knowledge tasks and model outputs for missing behaviors, weak reasoning, unsafe recommendations, departures from clinical guidelines, and plausible-sounding answers that do not meet clinical standards.
- Write detailed instruction specifications and gold-standard solutions for clinical problems.
- Create clinical tasks that reflect real-world medical practice.
- Design challenging benchmarks and evaluation sets to measure model improvement.
- Partner with researchers to develop medicine-specific capabilities and tools.
- Calibrate standards with researchers and adjacent-domain specialists, turning implicit clinical judgment into clear, teachable criteria.
Qualifications
- MD or DO from an accredited medical school, with completed residency training in a recognized specialty.
- At least 4 years of post-residency clinical practice. Residency and fellowship training do not count toward this requirement.
- Active, unrestricted medical license in at least one U.S. state and board certification in your specialty.
- Established expertise in a clinical specialty, such as internal medicine, oncology, radiology, emergency medicine, surgery, psychiatry, or a medical subspecialty.
- Senior clinical progression, such as Attending Physician, Medical Director, Division Chief, Associate or full Professor, or Chief Medical Officer, with meaningful ownership of clinical decisions.
- Hands-on professional use of large language models and the ability to distinguish sound clinical reasoning from convincing but incorrect answers.
- Excellent written communication and the ability to provide precise, structured feedback.
- Experience in utilization management, clinical informatics, or medical affairs is a plus.
Work Terms
- Full-time W-2 hourly employment, with an initial 6-month commitment and reliable availability for 40 hours per week.
- Hybrid role based in the Bay Area, California. You must live in the Bay Area and be available to work on-site with the client team multiple days per week when required.
- This is not a remote position. Candidates outside the Bay Area must relocate at their own expense before the engagement begins; relocation assistance is not provided.
- You will work within the client’s tools alongside internal research teams and receive client-issued accounts and equipment.
- The position is a structured, role-based placement within an enterprise team, not a freelance engagement.
Compensation
$70 to $110 per hour.
Application Process
Opportunities may be discovered through an online job platform. Employment, onboarding, payroll, benefits, and compliance are managed by the employer of record for the engagement.
Equal Opportunity
Equal employment opportunity is provided without discrimination based on any legally protected characteristic. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans throughout the application process.
$224k - $356.5k
...tapping into the unlimited potential of AI to define the next era of computing. An... ...Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful... ..., shaping the roadmap, and sharing best practices.Work alongside model training, inference...SuggestedFull time$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work... ...quality materials reasoning looks like in practice and ensure model outputs can withstand technical...SuggestedHourly payFull timeLive inRelocationRelocation package$300k - $320k
...role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be... ...partners and industry standards bodies to align our evaluation practices with emerging best practices in responsible AI development...SuggestedWork at officeHome officeVisa sponsorshipRelocation package$60 - $100 per hour
...Role Overview Help advance frontier AI models by bringing senior insurance and actuarial judgment to the evaluation of real-world insurance work. You will work directly... ...Create tasks that reflect real insurance practice, along with challenging benchmarks and evaluation...SuggestedHourly payFull timeLive inRelocationRelocation package$184.7k - $324.8k
Research Scientist / Engineer, Foundation Model Evaluation Cupertino, California, United States... ...Qualifications 3+ years of experience in AI model evaluation, NLP, or a related area... ...ability to translate research insights into practical implementations Strong experimental...SuggestedRelocation$238k - $302k
...in simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language Models... ...with software design principles, coding best practices, testing methodologies, and version control software...Full timeRemote work- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
- Apple Inc. is seeking a Research Scientist/Engineer to design evaluation systems for foundation models powering Apple products. You will work hands‑on across evaluation design, experimentation, and cross‑team collaboration to drive model improvement and product quality...
$50 - $75 per hour
A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in...Hourly payContract work- Job Description - Member of Technical Staff (Language Model Evaluations) Location: San Francisco (preferred), Sydney, Melbourne, Brisbane About... ...Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises...
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours- Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM...Weekday work
$350k
...Our first goal is to democratize frontier AI R&D across scientific disciplines. We... ...AI research company and training our own models end-to-end. Our work spans areas such as... ...looking for a research engineer to build the evaluation infrastructure that tells us whether our...- A leading AI company is seeking a legal professional for a contractor role focused on evaluating AI model outputs in legal contexts. Candidates must hold a Juris Doctor (J.D.) and have more than 3 years of experience in law. The role involves reviewing complex legal hypotheticals...For contractors10 hours per week
$400 per month
About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows...$218.5k - $288k
...Scientist specializing in Small Language Models and AI Training, you will lead research and... ...performance language models tailored for practical applications. You will work closely... ...language models.Design, implement, and evaluate model training experiments to improve performance...Work at officeFlexible hours3 days per week$65 - $105 per hour
...Help improve how frontier AI models reason about real-world life sciences research. In this... ...will apply deep scientific judgment to evaluate research tasks and model outputs, define... ...define tasks that reflect real research practice. Design challenging domain-specific evaluation...Hourly payFull timeLive inRelocationRelocation package$100 - $150 per hour
...Role Overview Help shape how next-generation AI models perform real financial work by providing deep, practical finance expertise to a GenAI research team. You will... ...depth: Design challenging finance tasks and evaluation sets, and collaborate with researchers to build...Hourly payFull timeLive inRelocationRelocation package$195.2k - $262.2k
...infrastructure for the global AI economy. We are building a full... ...and enterprises from data and model training through to production... ..., task environments, and evaluation sets for reasoning, coding, tool... ..., and failure analysis. Practical understanding of modern LLM behavior...Full timeTemporary workImmediate startRemote work$272k - $431.25k
...we’re generating it! Our world model team is pushing the boundaries of multimodal AI, robotics, and world foundation... ...Research Manager to lead world-model evaluation and benchmarking across NVIDIA’s... ...in our hiring and promotion practices) on the basis of race, religion,...Full time- ...providing independent assurance and evaluating the company's risk management... ...for data scientists and AI developers who will power our... ...in frameworks for auditing models, including criteria like robustness... ...knowledge in data analytics practices, machine learning, AI, and...
$95.68k - $164.32k
...transformer architectures, foundation models, and telemetry-driven AI to impact millions of ArcGIS users... ...infrastructureDesign and implement evaluation frameworks that measure model quality... ...documents, experiment reports, and best-practice guidance for model development and...Worldwide$400 per month
...About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering...- ...San Francisco is seeking an innovative Quality Engineer for their AI products. This role blends ops, strategy, and analytics to... ...in leading labs, and ensure user satisfaction through effective evaluation baselines. Competitive salary and benefits offered, with a focus...
$85 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our... ...Use frontier AI coding agents to complete and evaluate complex engineering tasks. Review model-generated mobile application code for correctness, quality...Contract workSummer workRemote work$184k - $287.5k
...is redefining what is possible with AI, and the Relational Foundation Model team is helping lead that... ...models: you will design, build, and evaluate novel Transformer and graph neural... ...including in our hiring and promotion practices) on the basis of race, religion, color...Full time$175k - $215k
...states. The mission of the Waymo AI Foundations team is to develop... ...demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you... ...field of study, or equivalent practical experience Proficiency in...Full timeRemote work$192k - $278k
Lead model releases for Search, evaluating DeepMind release applicants against strict quality bars to determine... ...as Search evolves into a fully AI-enabled product.Design and execute end... ...in a technical field, or equivalent practical experience. 8 years of experience in...Shift work$75 - $115 per hour
...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific... ...drug development reasoning looks like in practice. Key Responsibilities Review...Hourly payFull timeContract workLive inRelocationRelocation package- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes, and...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Practicing Clinician for AI Model Evaluation. Be the first to apply!





