Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Data Scientist, Health AI Evaluation & Datasets

Full-time

Innodata

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

Healthcare is one of the highest-stakes domains for generative AI. Clinical accuracy, patient safety, regulatory compliance, health equity, auditability, and workflow fit are the bar for shipping anything real. Innodata partners with foundation model labs, medical AI startups, payers, providers, pharma, and digital health companies building LLMs, multimodal systems, and AI agents for healthcare and life sciences.

As an Applied Data Scientist, Health AI Evaluation & Datasets , you own the design, measurement quality, and clinical validity of datasets used to train, fine-tune, and evaluate health-domain models. You bring clinical or biomedical fluency and data science rigor: you can read a clinical guideline, payer policy, medical literature artifact, or patient communication workflow; translate it into a measurable dataset and evaluation plan; and defend the methodology to sophisticated clinical, data science, and ML stakeholders.

You will work in a tight pod with a Technical Solutions Architect, Applied Research Scientist, AI/ML Research Engineer, and Language Data Scientists. Your role is to make sure the data, rubrics, review workflows, and measurement evidence are clinically realistic, statistically defensible, compliant, and useful for evaluation and post-training.

What You’ll Own:

  • Translate customer goals — such as improving differential diagnosis, evaluating a clinical note summarizer, testing a RAG-based medical literature assistant, or creating preference data for patient-facing chatbots — into dataset specifications, taxonomies, rubrics, sampling plans, and acceptance criteria.
  • Make multimodal health AI a core focus: design training and evaluation datasets across clinical text, medical images, waveforms, structured EHR data, claims, trial data, medical literature, patient communications, payer policies, drug information, and other clinical artifacts, as well as use cases such as clinical reasoning, medical QA, note summarization, medical coding, patient communication, utilization management, and literature synthesis.
  • Design evaluations for retrieval-augmented and source-grounded health AI systems, including evidence citation, faithfulness, contraindication handling, guideline adherence, source freshness, and failure modes caused by incomplete, conflicting, or stale context.
  • Define sampling strategies, label schemas, inter-annotator agreement targets, adjudication workflows, SME review patterns, and quality thresholds in partnership with Language Data Scientists, clinicians, biomedical experts, and quality teams.
  • Build statistical and ML checks that make healthcare datasets trustworthy: stratified sampling across specialties and patient subgroups, bias and representation analysis, leakage detection, distribution shift checks, uncertainty estimates, reliability metrics, and subgroup performance analysis.
  • Partner with Applied Research Scientists and AI/ML Research Engineers to instrument datasets into evaluation and post-training pipelines, including rubric-grounded LLM-as-judge prompts, regression suites, model comparison workflows, experiment tracking, and model-improvement feedback loops.
  • Evaluate health AI behavior beyond surface accuracy: calibration, hallucination on safety-critical content, refusal appropriateness, robustness under ambiguity, equity across patient subgroups, and safe handoff in agentic or workflow-integrated systems. Reason concretely about clinical workflow fit: where outputs enter care delivery, what evidence a clinician or reviewer would need to trust them, when uncertainty must be surfaced, and how patient-facing, clinician-facing, payer, pharma, and operational use cases differ in risk.
  • Own data quality from source intake through delivery, including de-identified clinical text, medical literature, synthetic cases, structured records, client policies, and knowledge bases, with attention to PHI/PII handling, provenance, audit trails, versioning, and compliance documentation.
  • Stay current on the health AI landscape — regulatory developments such as FDA guidance on AI/ML-enabled medical devices and EU AI Act health provisions , benchmark releases such as MedQA , MedMCQA , and HealthBench , and emerging clinical evaluation methodology.
  • Support customer discovery and proposal work by scoping dataset programs, sizing annotation and SME review effort, identifying regulatory or data-access constraints, and explaining methodology choices to client clinical and ML leadership.
  • Contribute to Innodata internal IP: reusable health-domain taxonomies, evaluation rubrics, golden datasets, clinical review playbooks, dataset quality checks, and methodology templates.

You’ll Thrive in This Role If You Have:

  • 5+ years of data science experience , including at least 2+ years with healthcare, clinical, biomedical, payer, provider, pharma, life sciences, or comparable regulated health data .
  • Working knowledge of healthcare data and standards: EHR structure , clinical documentation conventions, ICD-10 , CPT , SNOMED CT , LOINC , RxNorm , and at least passing familiarity with FHIR , HL7 , or equivalent interoperability concepts.
  • Hands-on experience designing ML datasets, not just consuming them: writing annotation guidelines, sizing cohorts, setting quality thresholds, designing QA checks, and shipping data that downstream teams can train or evaluate on.
  • Familiarity with LLM-based health AI workflows, including prompt design, rubric-based evaluation, retrieval-augmented generation, LLM-as-judge methods, model comparison, and the limitations of automated evaluation in clinical contexts.
  • Strong Python and SQL ; comfort with pandas , scikit-learn , statsmodels or equivalent tools; and working familiarity with modern LLM tooling such as Hugging Face , evaluation frameworks, prompt development tools, or model APIs.
  • Statistical literacy across sampling design, bias and fairness analysis, inter-annotator agreement metrics ( Cohen or Fleiss kappa , Krippendorff alpha ), confidence intervals, significance testing where appropriate, error analysis, and the ability to push back when a number is being over-interpreted.
  • Solid grasp of healthcare privacy, compliance, and governance: HIPAA , de-identification standards ( Safe Harbor and Expert Determination ), practical mechanics of working with PHI safely, auditability, access control, and documentation fit for high-stakes or regulated AI programs.
  • Ability to work credibly with clinicians, biomedical SMEs, research scientists, engineers, technical solutions teams, annotators, and customer stakeholders.
  • A bias toward clinical realism: you would rather build a smaller dataset that reflects what clinicians, reviewers, patients, or care teams actually see than a larger dataset that looks impressive on paper but fails in practice.
  • Degree in a relevant field such as biostatistics, epidemiology, computational biology, health informatics, computer science with a health focus, statistics, a clinical degree with quantitative training, or equivalent demonstrated experience.
  • Clinical credentials are not required, but candidates must be able to work credibly with clinicians, biomedical SMEs, and health AI customers; candidates with MD , RN , PharmD , MPH , PhD , or health informatics backgrounds are especially encouraged.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Applied Data Scientist, Health AI Evaluation & Datasets in Remote vacancy
  • $150k - $175k

    Role Description As an Applied Data Scientist, Health AI Evaluation & Datasets, you own the design, measurement quality, and clinical validity of datasets used to train, fine-tune, and evaluate health-domain models. You bring clinical or biomedical fluency and data science... 
    Suggested
    Full time

    Innodata Inc.

    Remote
    5 days ago
  •  ...INOD) is a global data engineering company...  ...Intelligence (AI) are inextricably...  ...providing the data, evaluation frameworks, and human...  ...compliance, health equity, auditability...  ...sciences. As an Applied Data Scientist, Health AI Evaluation & Datasets , you own the design... 
    Suggested
    Full time
    Shift work

    Innodata

    Remote
    19 days ago
  • $150k - $175k

    Role Description As an Applied Data Scientist, Financial AI Evaluation & Datasets, you own the design, measurement quality, and domain validity of the datasets used to train, fine-tune, evaluate, and monitor financial-domain LLMs, vision-language models, multimodal document... 
    Suggested
    Full time

    Innodata Inc.

    Remote
    5 days ago
  •  ...is one of the highest-stakes domains for generative AI. Numerical accuracy, regulatory compliance, model risk...  ..., and AI agents for financial workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you own the design, measurement quality, and... 
    Suggested
    Full time
    Shift work

    Innodata

    Remote
    a month ago
  •  ...is one of the highest-stakes domains for generative AI. Numerical accuracy, regulatory compliance, model risk...  ..., and AI agents for financial workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you own the design, measurement quality, and... 
    Suggested
    Full time
    Shift work

    Innodata

    Remote
    19 days ago
  • $154k - $200k

     ...looking for a Senior Data Scientist, Applied ML to design, build,...  ...model monitoring and evaluation, designing feedback loops...  ..., and pipeline health across the handoff points...  ...with cybersecurity datasets or domains: threat intelligence...  ...analytics and AI to accelerate... 
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Worldwide
    Visa sponsorship
    Flexible hours

    SpyCloud

    Austin, TX
    16 hours ago
  • $175k - $275k

     ...Product & Engineering - $175k - $275k Applied Data Scientist, LLM Evaluation Introduction At Driver, we’re...  ...the context layer for employees and AI agents alike to use in developing software...  ...metrics and build evaluation datasets. Establish what “good” looks like for... 
    Remote job
    Full time
    Flexible hours

    Driverai

    Austin, TX
    16 hours ago
  •  ...Learning Engineer, Data & Intelligence Products...  ...Growth, and Ajax Health. We're a high-growth AI and Data company scaling...  ...data assets by applying statistical and machine...  ...claims, and more datasets, to improve the coverage...  ...design, and model evaluation — and you know when... 
    Remote job
    Full time
    For contractors
    Work at office
    Work from home
    Home office
    Flexible hours

    Acuitymd

    Boston, MA
    16 hours ago
  • $98k - $176k

     ...dependents comprehensive health benefits and...  ...TARGET AS A SR DATA SCIENTIST - RECOMMENDATIONS...  ...Whether you join our Applied Data Sciences or Machine...  ..., Advanced AI, Search and Personalization...  ...designExperience evaluating models through...  ...on large-scale datasets and online experimentation... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Target

    Brooklyn Park, MN
    3 days ago
  • $127.1k - $236.1k

     ...world’s most complex health challenges and...  ...Principal Supply Chain Data Scientist, you will lead...  ...distribution operation.Apply machine learning,...  ....Analyze large datasets and business...  ....Develop advanced AI-driven tools for improved...  ...analysis, evaluating how effectively warehouse... 
    Full time
    Local area
    Relocation

    Genentech

    Louisville, KY
    16 hours ago
  • $130k - $160k

    At Nestlé Health Science, we believe that nutrition...  ...for a Senior Data Scientist to develop and...  ...advanced analytics and AI solutions that...  ...large and complex datasets to identify patterns...  ...track, evaluate, and refine deployed...  ...Learning Engineer, or Applied Scientist, specializing... 
    Remote work

    Nestlé

    Bridgewater, NJ
    16 hours ago
  • $129.57k - $194.35k

     ...leading not-for-profit health and well-being...  ...Audit.The Sr Data Scientist is responsible for...  ...high-dimensional datasets. This role leverages...  ...statistical techniques, and AI-powered...  ...analytical questions. Apply advanced machine learning...  ...in model evaluation, tuning, performance... 
    Full time
    Work experience placement
    Work at office
    Work from home

    Point32Health

    Canton, MA
    3 days ago
  • $96k - $107k

     ...UTAcute & Payer - Data Science /Full-...  .... As a leading health tech company...  ...post-acute care dataset and a Marketplace...  ...accelerated by AI to create...  ...Associate Data Scientist to join our Corporate...  ...KPIs, outcomes evaluation, or ROIRespond...  ....When you apply for a position,... 
    Full time
    Work at office
    Remote work
    Flexible hours

    PointClickCare

    Salt Lake City, UT
    16 hours ago
  • $8k

     ...fully FUNDED opportunity for a Senior Applied AI Engineer – Data Science & Analytics on our largest...  ...analysis, and machine learning on mission datasets - Develop pattern-of-life, anomaly...  ...into scalable AI applications - Evaluate emerging AI and machine learning... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Immediate start
    Flexible hours

    Visionist, Inc.

    Remote
    16 hours ago
  • $160k - $174k

     ...culture? Start your Voyage -Apply NowGet to Know the...  ...highly motivated Senior AI Engineer to design,...  ...applications on the Snowflake AI Data Cloud. This role...  ...unstructured enterprise datasets to identify AI...  ...retrieval tuning, and evaluation.AI Engineering & Agentic... 
    Full time
    Part time
    Work experience placement
    Local area
    Remote work
    Flexible hours

    Benefitfocus

    New York, NY
    3 days ago
  • $148k - $184k

    Role Description As a Data Scientist on our team, you will help define how we measure and...  .... You'll work at the intersection of applied AI evaluation, analytics engineering, and clinical...  ...reporting needs with accurate, well-tested datasets. ~For senior applicants: mentor... 
    Full time

    Paradigm Health

    Remote
    2 days ago
  •  ...for a product-minded Applied Data Scientist or Machine Learning Engineer...  ...both worlds to bring AI to life. WHAT YOU’...  ...& Strategy Evaluation & Experimentation:...  ...outcomes. Data Health & Feedback Loops: Collaborate...  ..., imperfect product datasets is essential.... 
    Full time
    Casual work
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    WorkWave

    Remote
    a month ago
  • $160k - $195k

     ...of employees’ financial health at every stage of life,...  ...building practical AI-powered solutions that...  ...looking for a Senior Applied AI Engineer to help design...  ...Develop solutions for evaluating and improving LLM output...  ...vector databases or AI data pipelines Experience... 
    Remote job
    Full time
    Local area

    Brightplan

    United States
    16 hours ago
  • $125k - $155k

     ...of employees’ financial health at every stage of life,...  ...building practical AI-powered solutions that...  ...We are looking for an Applied AI Engineer to help build...  ...design, and structured data work to deliver user-facing...  ...with prompt evaluation or LLM testing approaches... 
    Remote job
    Full time
    Local area

    Brightplan

    United States
    16 hours ago
  • Product Data Scientist (Product Analytics / ML) Full...  ...working. Peer AI produces far richer...  ...metrics for product health, activation, adoption...  ...analytical datasets and connect product...  ...to-Have Experience applying ML to anomaly detection...  ...AI/agent systems, evaluation data, or human... 
    Full time
    Remote work
    Home office

    Getpeer

    San Francisco, CA
    2 days ago
  •  ...developing the metrics and evaluation frameworks that ensure...  ...and simulated driving data, enabling data‑driven...  ...a skilled Data Scientist to join our team and play...  ...involve analyzing large datasets, applying statistical methods,...  ...please email ****@*****.***.ai. #J-18808-Ljbffr... 
    Remote work
    Relocation

    Avride

    Austin, TX
    16 hours ago
  •  ...Broward Health Corporate ISC Shift: Shift 1...  ...and maintain scalable data pipelines for the ingestion...  ...transformation of large complex datasets used in AI model training and...  ...by iteratively evaluating results and adjusting...  ...*Bonus Exclusions may apply in accordance with policy... 
    Shift work

    BHC ISC

    Fort Lauderdale, FL
    3 days ago
  •  ...) by giving generative AI models a better understanding...  ...We’re looking for a Applied AI/Machine Learning...  ..., character, and story data. Prototype and apply...  ...principles, algorithms, and evaluation metrics. Strong...  ...concentration in the Bay Area. ~ Health insurance for you and... 
    Remote job
    Full time
    Flexible hours

    Intangible

    United States
    16 hours ago
  • About Pivotal Health Pivotal Health is the leading technology...  ...combines software, data, and service into a seamlessly integrated, AI-driven platform that...  ...organization. As a Senior Applied AI/ML Engineer, you will...  ...through deployment, evaluation, monitoring, and continuous... 
    Permanent employment
    Full time
    Remote work
    Visa sponsorship
    Flexible hours

    Pivotal Health

    Los Angeles, CA
    16 hours ago
  • $158k - $176k

     ...people. As a leading health tech company that’s...  ...post-acute care dataset and a Marketplace of...  ...and accelerated by AI to create meaningful...  ...and Trino for big data processing. What...  ...principles and how they apply to large-scale data...  ...information to evaluate your candidacy for... 
    Remote job
    Full time
    Work at office
    Flexible hours

    Pointclickcare

    Remote
    16 hours ago
  • $100 per hour

     ...Role Overview Apply deep domain expertise to train and evaluate next-generation AI systems by producing, refining, and validating high-quality, domain-informed data. This part-time contractor role focuses...  ...evaluation. Interpret complex datasets or findings and summarize... 
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    Indiana
    19 days ago
  •  ...globally. Combining built-in AI and powerful workflow...  ...Machine Learning Engineer or Applied Data Scientist who can take an idea from concept...  ...customer, and operational data Evaluate ambiguous ideas quickly and...  ...plan (with no exclusions) Health — Health, dental, and vision... 
    Full time
    Flexible hours

    Gitkraken

    Remote
    16 hours ago
  •  ...building the leading AI-native platform for...  ...Our platform helps health plans and...  ...looking for a Senior Applied AI Engineer to design...  ...behavior should be evaluated, where deterministic...  ...structured and unstructured data, use internal tools...  ...representative datasets, evaluation... 
    Full time
    Work experience placement
    Relocation

    Abby Care

    San Francisco, CA
    1 day ago
  •  ...Senior Applied AI Engineer Roger is an AI platform that frees home health clinicians from paperwork so they can...  ...structure from unstructured data, validate their own...  ...a vast and growing dataset, and continuously improve...  ..., fine-tuning, or evaluating LLMs and open source models... 
    Remote work
    Work from home
    Flexible hours

    Roger Healthcare

    San Francisco, CA
    16 hours ago
  • $140k - $180k

     ...Mary West, West Health includes the nonprofit...  ...of seniors. Data/Data Science is an...  ...seeking a Senior Applied AI Engineer to join our...  ...patterns in large datasets, and recommend relevant...  ...instructions, and evaluation frameworks that...  ...with data scientists, analysts, visualization... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Monday to Friday
    Flexible hours

    West Health

    San Diego, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Data Scientist, Health AI Evaluation & Datasets. Be the first to apply!