Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Research Scientist, LLM Evaluation & Post-Training

Full-time

Innodata

Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.

This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.

The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.

What You’ll Own:

As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.

Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.

This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):

  • Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
  • Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
  • Develop and validate evaluation frameworks for LLM and multimodal systems, including:
    • benchmark/task design
    • scoring methods
    • judge/model-assisted evaluation
    • human evaluation protocols
    • robustness/stress testing
  • Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
  • Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
  • Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
  • Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
  • Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
  • Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
  • Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
  • Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
  • Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement

You’ll Thrive in This Role If You Have:

  • MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
  • 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
  • Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
  • Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
  • Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
  • Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
  • Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
  • Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
  • Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Applied Research Scientist, LLM Evaluation & Post-Training in Remote vacancy
  • $175k - $225k

    Role Description Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation... 
    Training
    Full time

    Innodata Inc.

    Remote
    7 days ago
  • $126.2k - $264.1k

     ...As our Principal Applied Scientist, you will play a key...  ...drive projects from research POC to production....  ...including data, model, training, and evaluation, employing best practices...  ...technologies in LLM and generative AI, such...  ...provided in this posting are specific to the... 
    Training
    Temporary work
    Remote work
    Flexible hours

    Ll Oefentherapie

    Montgomery, AL
    2 days ago
  • $100 - $120 per hour

     ...LLM Research Scientist (Pre-training & Post-Training) is a remote review track for evaluating AI outputs across research science reasoning, calculations, and research workflows....  ...Graduate-level training or equivalent applied experience in research science or a closely... 
    Training
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    23 hours ago
  •  ...power grids, buildings, trains, hospitals. Industrial...  ...major opportunity in applied AI, and one of the...  ...intersection of machine learning research, real world data, and...  .... As Applied Scientist, you work on the science...  ...control, planning, or evaluation. Take problems from ambiguous... 
    Training
    Local area

    Siemens Mobility

    New York, NY
    1 day ago
  •  ...The Role We're looking for an Applied Scientist who thrives at the intersection of applied research and real-world products. You'll...  ...in data, interaction, and evaluation with both creativity and engineering...  ...of data modeling for training and how curation decisions shape... 
    Training
    Flexible hours

    Adaption Labs

    San Francisco, CA
    2 days ago
  • $167.1k - $226.1k

     ...and Catalog Systems (ASCS) – Applied Scientist At Amazon Selection and...  ...groceries to digital content. The research challenges are immense....  ...data with rigorous evaluation frameworks Define research...  ...Qualifications Experience with training and deploying machine learning... 
    Training
    Worldwide

    Amazon

    Boston, MA
    2 days ago
  • $80 - $150 per hour

     ...future of AI systems by applying your physics expertise. As...  ...play a critical role in evaluating and enhancing the training of next-generation AI models...  ...arguments generated by researchers or AI platforms. Detect...  ..., or senior research scientist. Recent (last ~5 years... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Indiana
    15 days ago
  • $75 per hour

     ...Role Overview Physics experts apply advanced physics training to evaluate AI-generated scientific content and provide detailed feedback that improves...  ...full-time or tenure-track role, suitable to combine with research, teaching, postdoctoral work, or industry employment.... 
    Training
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $88.6k - $150.8k

    Join to apply for the Research Scientist role at The Henry M. Jackson Foundation for...  ...clinical tool development and training, cognitive monitoring and...  ...on employee performance evaluations. Qualifications Education...  ...Learning Scientist, NLP/LLM San Francisco, CA $140,000... 
    Training
    Full time
    Contract work
    For contractors
    Fixed term contract
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours

    The Henry M. Jackson Foundation for the Advancement of Milit...

    California, MO
    9 hours ago
  •  ...skilled and driven Senior Applied Scientist with 7+ years of industry experience...  .... Responsibilities ~Research and engineer NLP solutions...  ...and implement algorithms, train state of the art large language models (LLM) on large data, and evaluate their performance. ~... 
    Training
    Full time

    Black Ore

    Remote
    9 hours ago
  • $60 - $90 per hour

     ...collaborate closely with researchers to convert findings into robust evaluation benchmarks. Key...  ...familiarity with LLM capabilities,...  ...in AI model training, model evaluation,...  ...disabilities. How to Apply If you found this...  ...or the original posting. The employer-of-record... 
    Training
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    2 days ago
  •  ...record of exceptional research or engineering achievement...  ...… … you should apply for this role! About...  ...learning, and small custom post-trained models (SFT and RLVR)...  ...AI Research Scientist to join our small team...  ...from data generation to evaluation to product integration... 
    Training
    Full time
    Relocation package

    P-1 AI

    San Francisco, CA
    4 days ago
  • $40 per hour

     ...We are looking for an Applied Mathematician to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level of... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work

    DataAnnotation

    Brooklyn, NY
    3 days ago
  • $150k - $230k

     ...Engineer to drive the post-training of our large language...  ...large GPU clusters , applying distributed-training...  ...Build and maintain evaluation and reward/verifier pipelines...  ...with post-training research and turn promising...  ...Requirements Hands-on LLM post-training... 
    Training
    Full time
    Local area
    Work from home

    News Break

    Remote
    9 hours ago
  • $75 per hour

     ...Role Overview AI & Machine Learning Researchers apply advanced computer science and machine learning research expertise to evaluate AI-generated responses and produce expert-level training data that improves AI understanding of contemporary CS and AI/ML research and... 
    Training
    Hourly pay
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    4 days ago
  • A leading data services company is seeking an Applied Mathematician to evaluate AI models by providing complex mathematical problems to chatbots and assessing their outputs for quality and performance. This role offers flexibility with fully remote work and allows you to... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    3 days ago
  •  ...frontier of machine learning research. Founded by leading...  ...systems, and advanced training methodologies. These...  ...large-scale pre-training, post-training, reinforcement...  ...and data curation. Evaluation, benchmarking, reward modeling...  ...used at scale. Why Apply You’ll gain access to opportunities... 
    Training
    Currently hiring

    Alexander Chapman

    California, MO
    4 days ago
  • $204k - $259k

     ...service and can also be applied to a range of vehicle...  ...with other research teams in Alphabet. AI...  ...learning, and robust evaluation. Role Summary In this...  ...report to a Principal Scientist. Responsibilities Participate...  ...World Model post‑training and evaluation Research... 
    Training
    Temporary work
    Remote work

    Waymo

    New York, NY
    4 days ago
  • $100k

     ...Research Scientist Shift Type: -- ; Education: Doctorate ; Location: PRE...  ...national presence in the area of applied science, with funded...  ...implementation of complex program evaluation strategies and the conduct...  ...and scientists provide training and technical assistance in... 
    Training
    Full time
    Temporary work
    Part time
    Remote work
    Flexible hours
    Shift work

    Pacific Institute for Research and Evaluation

    Santa Fe, NM
    2 days ago
  • $92.15k

     ...1/2026; Rank CR - Research Asst Professor; Working...  ...Title Research Scientist; Position Number D...  ..., research, and training unit for population...  ...to collect and evaluate all necessary data...  ...of experience in applied social science research...  ...Search Details Posting Close Date; Projected... 
    Training
    Permanent employment
    Fixed term contract
    Remote work
    Relocation

    Portland State University

    Portland, OR
    5 days ago
  • $175k - $215k

     ...hail service and can also be applied to a range of vehicle platforms...  ...will: Design, implement, and evaluate state‑of‑the‑art generative...  ...tested code to bring cutting‑edge research into production Partner with...  ..., experience, relevant training and education, and skill level... 
    Training
    Full time
    Internship
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $183.8k - $248.7k

     ...ads. Key Job Responsibilities Lead the research and development of ML models that personalize...  ...analyze large‑scale A/B experiments, applying causal inference techniques to measure...  ...model fundamentals—including architecture, training/inference lifecycles, and optimization... 
    Training
    Flexible hours

    Prime Video & Amazon MGM Studios

    New York, NY
    5 days ago
  •  ...is a non-profit AI research institute dedicated...  ...Conducting pre- and post-release adversarial evaluations of frontier models...  ...fine-tuning and post-training workflows to...  ...significant overlap between scientist and engineer roles....  ...is by applying directly via the application... 
    Training
    Full time
    Remote work
    Visa sponsorship

    Aisafety

    Berkeley, CA
    2 days ago
  • $183.8k - $248.7k

     ...the full ML lifecycle, from exploratory research and offline modeling to online experimentation...  ...analyze large-scale A/B experiments, applying causal inference techniques to measure...  ...fundamentals, including architecture, training/inference lifecycles, and optimization of... 
    Training
    Flexible hours
    Shift work

    Amazon

    Arlington, VA
    2 days ago
  •  ...The Role We’re looking for Applied Scientists to join Wayve Labs and help...  ..., we are a high‑conviction research team with the strategic patience...  .... Define and evolve Evaluation Frameworks and Benchmarks for...  ...transformers, MoE, large‑scale training). Generative world modeling... 
    Training
    Full time
    Work at office
    Work from home
    Visa sponsorship
    Relocation package
    Flexible hours

    Icehouseventures

    Sunnyvale, CA
    2 days ago
  • $69.76k

     ...Responsibilities: The Research Scientist I will work as part...  ...outcomes to monitor and evaluate the effects of risk...  ...specified in the job posting and must also be located...  ...educational or training opportunities. Retirement...  ...are welcome to apply. Work Location Expectations... 
    Training
    Contract work
    Part time
    Work at office
    Local area
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Health Research

    Albany, NY
    4 days ago
  • $117.3k - $175.9k

    # Battery Research ScientistOnsite |Mid Level|Full TimePosted...  ...) within PSL conducts applied research in energy...  ...for characterization, evaluation, and assessment in...  ...As a Battery Research Scientist (Senior Member of Technical...  ...which includes: training, motivating and directing... 
    Training
    Full time
    For contractors
    Immediate start
    Remote work
    Work visa
    Relocation package
    Flexible hours

    SwiftCruit

    El Segundo, CA
    5 days ago
  • $200k - $250k

     ...'ll have the compute power to train large models to solve this challenge. Your role as Applied Scientist is to build next-generation AI...  ...fall short. Your focus: Applied research from research to data to model...  ...and evals. Training (pre & post-training) and fine-tuning Large... 
    Training
    Remote work
    Flexible hours

    techire ai

    Seattle, WA
    2 days ago
  • $150k - $173k

     ...join a small but world-class Applied Research and AI team and work on...  ...work in practice. Simulation & Evaluation: Build simulation environments...  ...- and can be used to train, evaluate, and iterate on decision...  ...monitoring and retraining RL systems post-deployment, including... 
    Training
    Bi-weekly pay
    Shift work

    The Nuclear Company

    Washington DC
    2 days ago
  • $126k - $167k

     ...Research Scientist, Battlespace Awareness Anduril Industries is a defense...  ...possess an M.S. or Ph.D. in Applied or Computational Mathematics...  ...interview process in which we also evaluate practical experience and...  ..., education and/or training, critical skills, and/or business... 
    Training
    Full time
    For contractors
    Work experience placement
    For subcontractor
    Flexible hours

    Anduril-1

    Fort Collins, CO
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!