Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Research Scientist, LLM Evaluation & Post-Training

Full-time

Innodata

Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.

This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.

The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.

What You’ll Own:

As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.

Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.

This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):

  • Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
  • Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
  • Develop and validate evaluation frameworks for LLM and multimodal systems, including:
    • benchmark/task design
    • scoring methods
    • judge/model-assisted evaluation
    • human evaluation protocols
    • robustness/stress testing
  • Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
  • Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
  • Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
  • Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
  • Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
  • Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
  • Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
  • Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
  • Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement

You’ll Thrive in This Role If You Have:

  • MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
  • 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
  • Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
  • Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
  • Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
  • Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
  • Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
  • Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
  • Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
Vacancy posted 29 days ago
Similar jobs that could be interesting for youBased on the Applied Research Scientist, LLM Evaluation & Post-Training in Remote vacancy
  • $245k - $315k

     ...Applied Research Scientist, LLM Evaluation & Post-Training Remote - Canada Innodata is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial... 
    Training
    Remote work

    Innodata Inc.

    United States
    13 hours ago
  • Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross... 
    Suggested
    Remote job
    Hourly pay
    Flexible hours

    AIToolboard

    New York, NY
    1 day ago
  •  ...This role builds the evaluation and decision systems that...  ...own and extend our LLM and agent evaluation pipeline...  ...consumption. You will research and prototype adaptive...  ..., agentic systems, applied machine learning, or...  ...opportunities such as training and mentoring.... 
    Training
    Full time

    Bitdeer Technologies Group

    Austin, TX
    2 days ago
  •  ...business problems.We’re training and deploying...  ...is a team of researchers, engineers, designers...  ...us!Why this role?Evaluation is critical to making...  ...to measure LLM progress.As a Senior Research Scientist, Model Evaluation,...  ...encourage you to apply. We strive to create... 
    Training
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    13 hours ago
  • $100 - $120 per hour

     ...empirical machine learning research across computer vision and...  .... This role focuses on training, improving, evaluating, and deploying deep learning...  ...measures such as FID; and applying training-efficiency techniques...  ...-parameter generators. LLM post-training and behavioral... 
    Training
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    28 days ago
  • $100 - $120 per hour

     ...LLM Research Scientist — Pre-training, Computer Vision, Adversarial Robustness is a remote red-team track for...  ...this role matters Adversarial evaluation is how AuraOne hardens AI models before...  ...120 / hr Application process Apply through AuraOne's specialist intake... 
    Training
    For contractors
    Remote work
    10 hours per week

    AuraOne Human Data

    Remote
    20 days ago
  • $136k

     ...models, large language models (LLM), generative audio (music...  ...part of a close-knit team of applied scientists and product managers who are...  ..., and bring cutting edge research to raise the bar within the...  ...Experience with LLM model training and tuning PREFERRED QUALIFICATIONS... 
    Training

    Amazon.com, Inc.

    Seattle, WA
    13 hours ago
  •  ...you?As a Principal Applied Scientist, you will lead the architecture, research, and...  ...systems, covering LLM fine-tuning, multimodal...  ...services for model training, fine-tuning, large...  ...and metrics for evaluating agentic systems,...  ...at the top of the posted range based on the... 
    Training
    Work at office
    Immediate start
    Remote work

    UiPath

    Bellevue, WA
    13 hours ago
  • $192k - $304.75k

     ...looking for a passionate scientist at the intersection...  .... As a Sr. Quantum Applied Research Scientist, you will help...  ...synthesis pipelines, post-trainable model...  ...will span synthetic training data generation, surrogate...  ..., fine-tuning, and evaluation.Strong background in... 
    Training
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    13 hours ago
  • $224k - $356.5k

     ...Curator team is seeking a Senior Applied Research Scientist with experience researching,...  ...for foundation-model training, including capabilities that...  ...in the training of Nemotron LLM models which lead the field...  ...effectively.Writing papers, blog posts, documentation and training... 
    Training
    Full time
    Work at office
    Remote work
    Flexible hours

    Nvidia

    Texas
    13 hours ago
  • $142.8k - $274.8k

     ...than 25%Profession: Research, Applied, & Data SciencesDiscipline...  ...a Principal Applied Scientist, you’ll lead the...  ...productionize large‑scale DNN/LLM‑enhanced recommenders...  ...business goals.Own evaluation and experimentation....  ...stores, distributed training/inference, and... 
    Training
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    10 hours ago
  • $183.8k - $248.7k

    As a Senior Applied Scientist specializing in lead scoring and...  ...teams to turn novel research into scalable,...  ...preprocessing, distributed training, model optimization,...  ...Define offline and online evaluation frameworks; establish...  ...Techniques**: LLM fine-tuning, federated... 
    Training
    Flexible hours

    AmazonWebServices

    Seattle, WA
    3 days ago
  • $142.8k - $193.2k

     ...talented, and inventive Applied Scientist to help build industry-leading...  ...understanding, modern LLM architectures, LLM evaluation & tooling, and a passion...  ...assistants.* Fine-tune/post-train LLMs using techniques...  ...delivery of solutions from research to production, including... 
    Training
    Flexible hours

    Amazon

    Bellevue, WA
    1 day ago
  • $126.2k - $264.1k

     ...As our Principal Applied Scientist, you will play a key...  ...drive projects from research POC to production....  ...including data, model, training, and evaluation, employing best practices...  ...technologies in LLM and generative AI, such...  ...provided in this posting are specific to the... 
    Training
    Temporary work
    Remote work
    Flexible hours

    Oracle

    United States
    2 days ago
  •  ...Applied Scientist Caseware is one of Canada's original Fintech companies...  ...in experiments, has built evaluations for LLM-based systems, and wants...  ..., machine learning, research, or data-intensive engineering...  ...(classical ML, model training, or MLOps). Experience... 
    Training
    Remote work
    Home office

    CaseWare

    United States
    13 hours ago
  •  ...skilled and driven Senior Applied Scientist with 7+ years of industry experience...  .... Responsibilities ~Research and engineer NLP solutions...  ...and implement algorithms, train state of the art large language models (LLM) on large data, and evaluate their performance. ~... 
    Training
    Full time

    Black Ore

    Remote
    a month ago
  • $300k

     ...Machine Learning Scientist 4 - Generative Models, Evaluation based in United...  ...learning research role focused on...  ...role combines applied research, model...  ...evaluation. Build LLM-based...  ...through model training, experimentation...  ...one or more LLM post-training approaches... 
    Training
    Full time
    Remote work
    Flexible hours

    jobgether

    United States
    9 days ago
  •  ...autonomously. As a Principal Applied Scientist, youll help invent...  .... You will lead the research and development of...  ...workflows, and post-training techniques to build autonomous...  ..., training and evaluating large models, developing...  ...model training LLM post-training Reinforcement... 
    Training
    Work at office
    Immediate start
    Remote work

    UiPath, Inc.

    Brooklyn, NY
    1 day ago
  • $100 - $120 per hour

     ...-scoped, empirical ML research projects that advance...  ...model development, from training models from scratch to...  ...and TRADES. Evaluating robust accuracy under...  ...metrics like FID. Applying training-efficiency techniques...  ...parameter counts. LLM Post-Training and Behavioral... 
    Training
    Hourly pay
    Temporary work
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  • $100 - $120 per hour

     ...LLM Research Scientist (Pre-training & Post-Training) is a remote review track for evaluating AI outputs across research science reasoning, calculations, and research workflows....  ...Graduate-level training or equivalent applied experience in research science or a closely... 
    Training
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  • $173.1k - $303k

     ...leads groundbreaking research, engineering, building...  ...data curation, training, and evaluation. Our goal is to consistently...  ...and creativity to apply existing methods and...  ...developers, applied research scientists, product managers and...  ...and developing LLM based features Experience... 
    Training
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Victrays

    Santa Clara, CA
    2 days ago
  • $193.93k - $352.29k

     ...output can be trusted — evaluation, verification, and the...  ...standard of proof we apply to the vehicle. Our...  ...of every engineer and researcher at Nuro by 100x. Not a...  ...that means nothing. Post-train models on data nobody...  ...Direct experience with LLM agent systems — building... 
    Training
    Full time

    Nuro

    California
    22 days ago
  •  ...estimator does. As a Computer Vision Applied Research Scientist at Boon, you will own end-to-end...  ...Research & Architecture ~Design and evaluate novel multi-stage vision architectures...  ..., fusion strategies, loss functions, training regimes. ~Run rigorous experiments... 
    Training
    Full time

    Boon Technologies, Inc.

    Remote
    22 days ago
  •  ...sits within the applied science...  ...~8+ years of post-degree industry...  ...production — not research-only experience...  ...unstructured text ~LLM-based information...  ..., and post-training ~Knowledge distillation...  ...~End-to-end evaluation framework...  ...Mentor applied scientists and ML practitioners... 
    Training
    Full time
    Flexible hours

    Thomson Reuters

    Remote
    17 days ago
  • $139.1k - $231.9k

     ...clinical sciences and applied AI. The strongest candidate...  ...with clinical scientists, clinical operations,...  ...scientific requirements and evaluation criteria that make...  ...the responsible use of LLM-enabled tools. BASIC...  ...; building or training models is not required... 
    Training
    Permanent employment
    Full time
    H1b
    Local area
    Visa sponsorship
    Work visa
    Relocation package
    2 days per week

    Pfizer

    Cambridge, MA
    1 day ago
  • $142.8k - $193.2k

     ...Language Models (LLMs): a new LLM stack that already powers...  ...across Amazon.We are hiring an Applied Scientist to push the science behind this...  ...model lifecycle, from mid-training reasoning models on shopping...  ...and products.- Mid-train and post-train large language models on... 
    Training
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $119.8k - $234.7k

     ...than 25%Profession: Research, Applied, & Data SciencesDiscipline...  ...through clicks, post-click engagement, and...  ...outcomes. We design and train transformer-based...  ...rigorous offline and online evaluation. We build...  ...dynamics.Engineers and scientists on the team work at the... 
    Training
    Ongoing contract
    Work at office
    Local area
    Shift work

    Microsoft

    Redmond, WA
    1 day ago
  • $142.8k - $193.2k

     ...universally applicable signals and algorithms for training machine-learned ranking models. The...  ...solutions to improve search ranking.* Evaluate the proposed solutions via offline...  ...information. If the country/region you’re applying in isn’t listed, please contact your Recruiting... 
    Training
    Local area
    Worldwide
    Flexible hours

    Amazon

    Palo Alto, CA
    3 days ago
  • $117.13k - $135k

     ...Research Scientist Center for Effective Organization USC Marshall...  ...itself from other applied research centers and...  ...that assessment, and evaluating the effectiveness of...  ...Proficiency in use of LLM tools Expertise in organization...  ..., education/training, key skills, internal... 
    Training
    Work experience placement
    Local area

    University of Southern California

    Los Angeles, CA
    13 hours ago
  • $141.5k - $196k

     ...cases. The Role: As an Applied Scientist at Upstart, you will improve...  ...analysis, machine learning research, experimentation, and...  ...marketing outcomes. Develop and evaluate machine learning models,...  ..., and relevant education or training. Your recruiter can share more... 
    Training
    Summer work
    Currently hiring
    Local area
    Remote work
    Work from home

    UpStart

    United States
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!