Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Research Scientist, LLM Evaluation & Post-Training

Full-time

Innodata

Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.

This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.

The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.

What You’ll Own:

As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.

Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.

This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):

  • Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
  • Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
  • Develop and validate evaluation frameworks for LLM and multimodal systems, including:
    • benchmark/task design
    • scoring methods
    • judge/model-assisted evaluation
    • human evaluation protocols
    • robustness/stress testing
  • Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
  • Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
  • Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
  • Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
  • Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
  • Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
  • Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
  • Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
  • Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement

You’ll Thrive in This Role If You Have:

  • MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
  • 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
  • Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
  • Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
  • Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
  • Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
  • Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
  • Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
  • Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
Vacancy posted 13 days ago
Similar jobs that could be interesting for youBased on the Applied Research Scientist, LLM Evaluation & Post-Training in Remote vacancy
  •  ...Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how... 
    Training
    Full time

    Innodata

    Remote
    25 days ago
  •  ...Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how... 
    Training
    Full time

    Innodata

    Remote
    12 days ago
  • Research Scientist, LLM Evaluation & Post-Training page is loaded## Research Scientist, LLM Evaluation & Post-Traininglocations: Remote Work( USA)time type...  ...collaborative research role that sits at the intersection of applied ML research, enterprise AI product development, and... 
    Training
    Full time
    Remote work

    Centific Global Solutions, Inc.

    Seattle, WA
    3 days ago
  • Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross... 
    Suggested
    Remote job
    Hourly pay
    Flexible hours

    AIToolboard

    New York, NY
    5 days ago
  •  ...business problems.We’re training and deploying...  ...is a team of researchers, engineers, designers...  ...us!Why this role?Evaluation is critical to making...  ...to measure LLM progress.As a Senior Research Scientist, Model Evaluation,...  ...encourage you to apply. We strive to create... 
    Training
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    4 days ago
  • $100 - $120 per hour

     ...empirical machine learning research across computer vision and...  .... This role focuses on training, improving, evaluating, and deploying deep learning...  ...measures such as FID; and applying training-efficiency techniques...  ...-parameter generators. LLM post-training and behavioral... 
    Training
    Hourly pay
    Remote work
    Flexible hours

    SaidGig

    United States
    12 days ago
  • $136k

     ...models, large language models (LLM), generative audio (music...  ...part of a close-knit team of applied scientists and product managers who are...  ..., and bring cutting edge research to raise the bar within the...  ...Experience with LLM model training and tuning PREFERRED QUALIFICATIONS... 
    Training

    Amazon.com, Inc.

    Seattle, WA
    3 days ago
  • $50 per hour

     ...Chemistry specialists apply their subject-matter knowledge to create and evaluate prompts, review model outputs...  .... Evaluate and rate LLM responses for accuracy,...  ...chemistry contexts. Research chemistry topics to...  ...with valid work or training authorization, for example... 
    Training
    Part time
    H1b
    Remote work
    Visa sponsorship
    10 hours per week
    Flexible hours

    SaidGig

    United States
    2 days ago
  • $50 per hour

     ...Overview Physics specialists apply advanced physics knowledge to evaluate and improve large...  ...outputs, and conducting research-driven analyses. This part...  ...role focuses on improving LLM performance across physics...  ...graduates with valid work or training authorization, for... 
    Training
    Part time
    H1b
    Remote work
    10 hours per week

    SaidGig

    United States
    1 day ago
  • $192k - $304.75k

     ...looking for a passionate scientist at the intersection...  .... As a Sr. Quantum Applied Research Scientist, you will help...  ...synthesis pipelines, post-trainable model...  ...will span synthetic training data generation, surrogate...  ..., fine-tuning, and evaluation.Strong background in... 
    Training
    Full time
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  •  ...you?As a Principal Applied Scientist, you will lead the architecture, research, and...  ...systems, covering LLM fine-tuning, multimodal...  ...services for model training, fine-tuning, large...  ...and metrics for evaluating agentic systems,...  ...at the top of the posted range based on the... 
    Training
    Work at office
    Immediate start
    Remote work

    UiPath

    Bellevue, WA
    4 days ago
  • $183.8k - $248.7k

    As a Senior Applied Scientist specializing in lead scoring and...  ...teams to turn novel research into scalable,...  ...preprocessing, distributed training, model optimization,...  ...Define offline and online evaluation frameworks; establish...  ...Techniques**: LLM fine-tuning, federated... 
    Training
    Flexible hours

    AmazonWebServices

    Seattle, WA
    2 days ago
  • $142.8k - $193.2k

     ...talented, and inventive Applied Scientist to help build industry-leading...  ...understanding, modern LLM architectures, LLM evaluation & tooling, and a passion...  ...assistants. * Fine-tune/post-train LLMs using techniques...  ...of solutions from research to production, including... 
    Training
    Flexible hours

    Amazon

    Bellevue, WA
    4 days ago
  • $119.8k - $234.7k

     ...ContributorTravel: Less than 25%Profession: Research, Applied, & Data SciencesDiscipline:...  ..., including the latest LLM models. Deploying robust and...  ...Experience with distributed training or inference for SLMs and...  ...implementation.  Experience evaluating SLMs, LLMs, or agentic... 
    Training
    Ongoing contract
    Work experience placement
    Work at office
    Local area

    Microsoft

    Redmond, WA
    2 days ago
  • $126.2k - $264.1k

     ...As our Principal Applied Scientist, you will play a key...  ...drive projects from research POC to production....  ...including data, model, training, and evaluation, employing best practices...  ...technologies in LLM and generative AI, such...  ...provided in this posting are specific to the... 
    Training
    Temporary work
    Remote work
    Flexible hours

    Oracle

    United States
    1 day ago
  •  ...Applied Scientist, Alexa International Tech Alexa International...  ...to scientific research and applied AI for multi...  ...language processing, modern LLM architectures, LLM evaluation & tooling, and a...  ...-Speech (S2S) model training and fine-tuning Fine-tune/post-train LLMs using techniques... 
    Training

    Amazon

    Seattle, WA
    1 day ago
  •  ...autonomously. As a Principal Applied Scientist, you'll help invent...  ...You will lead the research and development of...  ...workflows, and post-training techniques to build autonomous...  ..., training and evaluating large models, developing...  ...model training LLM post-training Reinforcement... 
    Training
    Full time
    Work at office
    Immediate start
    Remote work

    UiPath

    Bellevue, WA
    3 days ago
  •  ...skilled and driven Senior Applied Scientist with 7+ years of industry experience...  .... Responsibilities ~Research and engineer NLP solutions...  ...and implement algorithms, train state of the art large language models (LLM) on large data, and evaluate their performance. ~... 
    Training
    Full time

    Black Ore

    Remote
    4 hours ago
  •  ...estimator does. As a Computer Vision Applied Research Scientist at Boon, you will own end-to-end...  ...Research & Architecture ~Design and evaluate novel multi-stage vision architectures...  ..., fusion strategies, loss functions, training regimes. ~Run rigorous experiments... 
    Training
    Full time

    Boon Technologies, Inc.

    Remote
    5 days ago
  • $147.6k - $274.2k

     ...sits within the applied science...  ...~8+ years of post-degree industry...  ...production — not research-only experience...  ...unstructured text ~LLM-based information...  ..., and post-training ~Knowledge distillation...  ...~End-to-end evaluation framework...  ...Mentor applied scientists and ML practitioners... 
    Training
    Full time
    Flexible hours

    Thomson Reuters

    Remote
    6 hours ago
  • $100 - $120 per hour

     ...LLM Research Scientist (Pre-training & Computer Vision & Adversarial Robustness) is a remote red-team track...  ...Why this role matters Adversarial evaluation is how AuraOne hardens AI models before...  ...$120 / hr Application process Apply through AuraOne's specialist intake... 
    Training
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    4 days ago
  • $100 - $120 per hour

     ...-scoped, empirical ML research projects that advance...  ...model development, from training models from scratch to...  ...and TRADES. Evaluating robust accuracy under...  ...metrics like FID. Applying training-efficiency techniques...  ...parameter counts. LLM Post-Training and Behavioral... 
    Training
    Hourly pay
    Temporary work
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $100 - $120 per hour

     ...LLM Research Scientist (Pre-training & Post-Training) is a remote review track for evaluating AI outputs across research science reasoning, calculations, and research workflows....  ...Graduate-level training or equivalent applied experience in research science or a closely... 
    Training
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    19 days ago
  •  ...Overview Use your Biology training to design and solve...  ..., help define evaluation benchmarks across undergraduate...  ...with model researchers to probe and improve model...  ...explanations. Work with LLM researchers to align...  ...eligible and encouraged to apply. Applicants must... 
    Training
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $193.93k - $352.29k

     ...output can be trusted — evaluation, verification, and the...  ...standard of proof we apply to the vehicle. Our...  ...of every engineer and researcher at Nuro by 100x. Not a...  ...that means nothing. Post-train models on data nobody...  ...Direct experience with LLM agent systems — building... 
    Training
    Full time

    Nuro

    California
    6 days ago
  • $167.1k - $226.1k

     ...do. You will distill LLM quality judgments into...  ...into online objectives. Scientists on this team own the...  ...combination of problems: research-grade modeling with an...  ...online. That means training compact models against...  ...degree and 6+ years of applied research experience- Experience... 
    Training
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $137.7k - $275.4k

     ...Incubation team seeks a Research Scientist to advance AI-driven...  ...group of PhDs and applied scientists working at...  ...operating at the frontier of LLM posttraining, agentic...  ...architectures and post-training strategies that...  ...contexts.Developing evaluation and benchmarking frameworks... 
    Training
    Full time
    Work at office
    Remote work
    Worldwide

    Zoom

    Seattle, WA
    2 days ago
  • $142.8k - $193.2k

     ...your needs.Key job responsibilitiesAs an Applied Scientist on the team, you will help lead science...  ...developing state of the art thinking-LLM-based techniques to reason about customers...  ...cross-attentive LLM rankers ; training multi-objective ranking and optimization... 
    Training
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $119.8k - $234.7k

     ...than 25%Profession: Research, Applied, & Data SciencesDiscipline...  ...through clicks, post-click engagement, and...  ...outcomes. We design and train transformer-based...  ...rigorous offline and online evaluation. We build...  ...dynamics.Engineers and scientists on the team work at the... 
    Training
    Ongoing contract
    Work at office
    Local area
    Shift work

    Microsoft

    Redmond, WA
    5 days ago
  • $117.13k - $135k

    Research Scientist Center for Effective OrganizationUSC...  ...from other applied research centers...  ...assessment, and evaluating the effectiveness...  ...skillsProficiency in use of LLM toolsExpertise in...  ..., education/training, key skills, internal...  ...Job openings are posted for a minimum of... 
    Training
    Full time
    Work experience placement
    Local area

    The University of Southern California

    Los Angeles, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!