Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Research Scientist, LLM Evaluation & Post-Training

$175k - $225k
Full-time

Innodata Inc.

Role Description

Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.

This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.

The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.

What You’ll Own

  • Define the next generation of evaluation-driven model improvement workflows.
  • Study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes.
  • Design experiments that produce credible, actionable conclusions.
  • Design benchmark datasets, develop evaluation taxonomies and protocols, define metrics and scoring methodologies, analyze failure modes, and test how changes in evaluation setup affect downstream fine-tuning results.
  • Support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
  • Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement.
  • Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes.
  • Develop and validate evaluation frameworks for LLM and multimodal systems, including:
    • benchmark/task design
    • scoring methods
    • judge/model-assisted evaluation
    • human evaluation protocols
    • robustness/stress testing
  • Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations.
  • Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs.
  • Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign.
  • Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines.
  • Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs.
  • Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations.
  • Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets.
  • Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations.
  • Contribute to thought leadership and best practices in LLM evaluation, post-training, and GenAI quality measurement.

Qualifications

  • MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field (PhD strongly preferred).
  • 5+ years of relevant experience in applied research / research science in ML/AI, with substantial work in LLMs or foundation models.
  • Demonstrated experience with LLM evaluation, benchmarking, alignment, post-training, or model quality research.
  • Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems.
  • Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization).
  • Experience working with modern ML tooling/frameworks (e.g., PyTorch, Hugging Face, JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments.
  • Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability.
  • Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs.
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts.
  • Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly.

Requirements

  • The expected salary range for this position is $175,000 – $225,000 USD per year, based on experience, skills, and qualifications.

Company Description

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Applied Research Scientist, LLM Evaluation & Post-Training in Remote vacancy
  • Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation... 
    Training
    Full time

    Innodata

    Remote
    12 hours ago
  • $142.8k - $193.2k

     ...individual customers, building evaluation frameworks for model...  ...designing data‑driven guardrails for LLM‑generated content. The work...  ...architecture design and deep learning training and optimization and model...  ...If the country/region you’re applying in isn’t listed, please... 
    Training
    Local area
    Worldwide
    Flexible hours

    Amazon Science

    Seattle, WA
    1 day ago
  • $171.6k - $222.2k

     ...You will design, train, and optimize generative...  ...involves building LLM-based classifiers...  ...data engines, and post-generation content...  ...metrics, defining evaluation frameworks that...  ...-knit group of scientists and engineers who...  ...country/region you’re applying in isn’t listed,... 
    Training
    Local area
    Flexible hours

    Amazon

    Palo Alto, CA
    1 day ago
  • $167.1k - $226.1k

     ...scalable and reliable evaluation of state-of-the-art Conversational...  ..., and resourceful Applied Scientist in the field of Large...  ...with LLMs, including LLM‑as‑a‑Judge (LLMaaJ),...  ...expertise to set the research agenda for how we...  ...sources, and design, train, and maintain the evaluation... 
    Training
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    1 day ago
  • $126.2k - $264.1k

     ...As our Principal Applied Scientist, you will play a key...  ...drive projects from research POC to production....  ...including data, model, training, and evaluation, employing best practices...  ...technologies in LLM and generative AI, such...  ...provided in this posting are specific to the... 
    Training
    Temporary work
    Remote work
    Flexible hours

    Ll Oefentherapie

    Montgomery, AL
    3 days ago
  • $100 - $120 per hour

     ...LLM Research Scientist (Pre-training & Post-Training) is a remote review track for evaluating AI outputs across research science reasoning, calculations, and research workflows....  ...Graduate-level training or equivalent applied experience in research science or a closely... 
    Training
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    7 days ago
  • $167.1k - $226.1k

     ...responsibilities As an Applied Scientist on LPEX, you will be a technical...  ..., translating research advances into measurable business...  ...development, offline and online evaluation, and reliable production...  ...architecture design and deep learning training and optimization and model... 
    Training
    Worldwide
    Flexible hours
    Night shift

    Amazon.com Services LLC

    Seattle, WA
    7 hours ago
  • $167.1k - $226.1k

     ...innovative and customer-focused applied scientist to help us make the world's...  ...at the frontier of AI research, and rapidly implement and...  ...multimodal data with rigorous evaluation frameworks. Define research...  ...Experience with training and deploying machine learning... 
    Training
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $149k - $350k

     ...join us! We’re looking for applied scientists with a Machine Learning and...  ...fundamental and applied research in this area. You will be combining...  ...for AI systems Build evaluation systems to measure and...  ...HuggingFace etc ~ Experience training LLMs with Reinforcement... 
    Training
    Full time
    Temporary work
    Remote work
    Work from home

    Figma

    New York, NY
    1 day ago
  •  ...business problems. We’re training and deploying...  ...is a team of researchers, engineers, designers...  ...us! Why this role? Evaluation is critical to...  ...infrastructure to measure LLM progress. As a Senior Research Scientist, Model Evaluation,...  ...encourage you to apply. We strive to... 
    Training
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  •  ...The Role We're looking for an Applied Scientist who thrives at the intersection of applied research and real-world products. You'll...  ...in data, interaction, and evaluation with both creativity and engineering...  ...of data modeling for training and how curation decisions shape... 
    Training
    Flexible hours

    Adaption Labs

    San Francisco, CA
    3 days ago
  •  ...power grids, buildings, trains, hospitals. Industrial...  ...major opportunity in applied AI, and one of the...  ...intersection of machine learning research, real world data, and...  .... As Senior Applied Scientist, you lead the science...  ...control, planning, or evaluation Take problems from... 
    Training

    Siemens

    New York, NY
    5 days ago
  •  ...focused, healthcare‑only LLM platform enables...  ...voice. As a Senior Staff Research Scientist in Speech Technologies...  ...datasets, creating the training foundation that gives...  ...medical domain Train, evaluate, and optimize ASR models...  ...to match. Ready to Apply? If you've spent your... 
    Training
    Work at office
    Remote work

    AI Chopping Block, Inc.

    Bellevue, WA
    18 hours ago
  • $80 - $110 per hour

     ...expertise to the development and evaluation of next-generation AI systems that must reason about pure and applied mathematics at the level of a working research mathematician. You will help ensure...  ...research-level mathematics problems to train and evaluate frontier AI models.... 
    Training
    Hourly pay
    Part time
    Immediate start
    Remote work

    SaidGig

    United States
    1 day ago
  • $80 - $110 per hour

     ...Contribute frontier research expertise to the development and evaluation of AI systems that...  ...supports model training and assessment across...  ...member, principal scientist, or industry...  ...are encouraged to apply, the depth of research...  ...application through the posting you are viewing,... 
    Training
    Hourly pay
    Part time
    Immediate start
    Remote work

    SaidGig

    United States
    1 day ago
  •  ...skilled and driven Senior Applied Scientist with 7+ years of industry experience...  .... Responsibilities ~Research and engineer NLP solutions...  ...and implement algorithms, train state of the art large language models (LLM) on large data, and evaluate their performance. ~... 
    Training
    Full time

    Black Ore

    Remote
    17 hours ago
  • $88.6k - $150.8k

    Join to apply for the Research Scientist role at The Henry M. Jackson Foundation for...  ...clinical tool development and training, cognitive monitoring and...  ...on employee performance evaluations. Qualifications Education...  ...Learning Scientist, NLP/LLM San Francisco, CA $140,000... 
    Training
    Full time
    Contract work
    For contractors
    Fixed term contract
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours

    The Henry M. Jackson Foundation for the Advancement of Milit...

    California, MO
    1 day ago
  • $40 per hour

     ...We are looking for an Applied Mathematician to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level of... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work

    DataAnnotation

    Brooklyn, NY
    4 days ago
  • $150k - $173k

     ...join a small but world-class Applied Research and AI team and work on...  ...in practice. * Simulation & Evaluation: Build simulation environments...  ...operations — and can be used to train, evaluate, and iterate on...  ...and retraining RL systems post-deployment, including... 
    Training
    Bi-weekly pay
    Full time
    Shift work

    The Nuclear Company

    Washington DC
    10 hours ago
  • $150k - $230k

     ...Engineer to drive the post-training of our large language...  ...large GPU clusters , applying distributed-training...  ...Build and maintain evaluation and reward/verifier pipelines...  ...with post-training research and turn promising...  ...Requirements Hands-on LLM post-training... 
    Training
    Full time
    Local area
    Work from home

    News Break

    Remote
    17 hours ago
  • $200k - $270k

     ...looking for a Staff Research Scientist to join Cognitiv’...  ...the AdTech and LLM landscape well...  ...embeddings, and applied deep learning to...  ...including large‑scale training, fine‑tuning, and...  ...of a launch—blog posts, talks, demos, or...  ...candidate evaluation or hiring decisions... 
    Training
    Work at office
    Remote work
    Work from home

    Cognitiv Corp

    Bellevue, WA
    1 day ago
  • $75 per hour

     ...Role Overview AI & Machine Learning Researchers evaluate AI-generated content and create expert-level training data to improve AI systems'' understanding of advanced computer...  ...(cs), and the ability to interpret and apply findings from recent preprints. Strong written... 
    Training
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • A leading data services company is seeking an Applied Mathematician to evaluate AI models by providing complex mathematical problems to chatbots and assessing their outputs for quality and performance. This role offers flexibility with fully remote work and allows you to... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    4 days ago
  • $311.85k - $370k

     ...Wayve AI Platform Scientist Or Engineer Founded in...  ...we curate, enrich, and evaluate the real-world driving...  ...are hiring at either Applied Scientist or Machine Learning...  ...into high-signal training data through (semi-)...  ...record of taking ML from research into production... 
    Training
    Full time
    Work at office
    Work from home

    Wayve

    Sunnyvale, CA
    4 days ago
  •  ...growth. If you are ready to apply your skills to the...  ...robotics stack. We're training state-of-the-art AI models...  ...As a Machine Learning Research Engineer, you will work...  ...to model training, evaluation, and on-robot deployment...  ...model training pipelines (LLM/VLM/VLA) At... 
    Training

    Sunday

    Redwood City, CA
    4 days ago
  • $200k - $250k

     ...'ll have the compute power to train large models to solve this challenge. Your role as Applied Scientist is to build next-generation AI...  ...fall short. Your focus: Applied research from research to data to model...  ...and evals. Training (pre & post-training) and fine-tuning Large... 
    Training
    Remote work
    Flexible hours

    techire ai

    Seattle, WA
    3 days ago
  • $183.8k - $248.7k

     ...the full ML lifecycle, from exploratory research and offline modeling to online experimentation...  ...analyze large-scale A/B experiments, applying causal inference techniques to measure...  ...fundamentals, including architecture, training/inference lifecycles, and optimization of... 
    Training
    Flexible hours
    Shift work

    Amazon

    New York, NY
    3 days ago
  • $30 per hour

    A leading AI training firm is seeking an Applied Physics Research Scientist to evaluate AI chatbots on physics-related challenges. This is a remote position ideal for experts with a strong grasp of classical mechanics and related fields. Applicants must possess fluency... 
    Training
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    1 day ago
  • $125k - $225k

     ...FutureSearch is looking for exceptional Research Scientists to evaluate and improve state-of-the-art forecasting and agentic LLM web research. We are an elite team of engineers...  ...frontier labs on research, evaluation, and training. You are a talented researcher with... 
    Training
    Remote work
    Flexible hours

    Future Research Corp

    San Francisco, CA
    1 day ago
  • $290.25k

     ...Machine Learning Scientists to join our AI team...  ...AI applications (LLM and Computer Vision...  ...key member of our research and development efforts...  ...architectures. Evaluate the performance of...  ...architectures, train/evaluate/tune models...  ...less likely to apply to jobs unless they... 
    Training
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!