Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Research Scientist, LLM Evaluation & Post-Training

Full-time

Innodata

Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.

This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.

The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.

What You’ll Own:

As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.

Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.

This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):

  • Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
  • Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
  • Develop and validate evaluation frameworks for LLM and multimodal systems, including:
    • benchmark/task design
    • scoring methods
    • judge/model-assisted evaluation
    • human evaluation protocols
    • robustness/stress testing
  • Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
  • Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
  • Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
  • Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
  • Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
  • Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
  • Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
  • Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
  • Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement

You’ll Thrive in This Role If You Have:

  • MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
  • 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
  • Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
  • Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
  • Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
  • Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
  • Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
  • Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
  • Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
Vacancy posted 24 days ago
Similar jobs that could be interesting for youBased on the Applied Research Scientist, LLM Evaluation & Post-Training in Canada vacancy
  • Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation... 
    Training
    Full time

    Innodata

    Canada
    12 days ago
  •  ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you...  .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client... 
    Training
    Full time

    Innodata

    Canada
    24 days ago
  •  ...Role We’re looking for Applied Scientists to join Wayve Labs and help...  ..., We Are a High-conviction Research Team With The Strategic Patience...  ...Define and evolve Evaluation Frameworks and Benchmarks for...  ...transformers, MoE, large-scale training) ~ Generative world... 
    Training
    Full time
    Work at office
    Work from home
    Visa sponsorship
    Relocation package
    Flexible hours

    Wayve

    Canada
    2 days ago
  •  ...providing the data, evaluation frameworks, and...  .... As an Applied Data Scientist, Health AI Evaluation...  ...datasets used to train, fine-tune, and evaluate...  ..., Applied Research Scientist, AI/ML...  ...for evaluation and post-training. What...  ...rubric-grounded LLM-as-judge prompts,... 
    Training
    Full time
    Shift work

    Innodata

    Canada
    12 days ago
  •  ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will...  .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client... 
    Training
    Full time

    Innodata

    Canada
    12 days ago
  • $80 - $90 per hour

     ...Role Overview Apply advanced biology expertise to help train and evaluate next-generation AI systems. You will provide scientifically rigorous, real-world input...  ...Experience analyzing scientific data, evaluating research critically, and synthesizing complex material.... 
    Training
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    20 days ago
  •  ...workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you...  ...the datasets used to train, fine-tune,...  ...engagement), an Applied Research Scientist (shapes...  ...for evaluation and post-training. What You...  ...— rubric-grounded LLM-as-judge prompts, regression... 
    Training
    Full time
    Shift work

    Innodata

    Canada
    24 days ago
  • $80 - $100 per hour

     ...Role Overview Apply deep computational physics knowledge to produce detailed, real...  ...technical feedback that improve model training and evaluation. Communicate findings accurately...  ...record of technical problem solving or research. Proficiency in Python and comfort... 
    Training
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    9 days ago
  • $80 - $160 per hour

     ...condensed matter and magnetism analysis to a research program focused on magnetic space...  ...high quality, real-world content that trains and evaluates next generation AI models. No prior...  ...knowledge of neutron scattering techniques as applied to magnetic structure determination.... 
    Training
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    15 days ago
  • $20 - $40 per hour

     ...Role Overview Apply your biology expertise to improve...  ...; your academic and research background is the key...  ...content. Author, evaluate, and edit laboratory reports...  ...materials used for AI training and validation....  ...sponsorship details are provided in this posting.... 
    Training
    Hourly pay
    Contract work
    For contractors
    Remote work
    Visa sponsorship

    SaidGig

    Canada
    a month ago
  •  ...Role Overview Apply your biochemistry expertise...  ...contractor role you will evaluate biochemical data,...  ...biochemical datasets, research findings, and laboratory...  ...biochemistry practice for AI training purposes. Assess...  ...through this posting. Include your profile... 
    Training
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    20 days ago
  • $80 - $150 per hour

     ...Role Overview Apply deep physics expertise to evaluate, verify, and improve scientific...  ...reasoning used to train next-generation AI...  ...produced by researchers or AI systems. Detect...  ...or senior research scientist. Recent...  ...qualification. This posting is currently closed... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    a month ago
  •  ...Role Overview Design and evaluate high-quality datasets and evaluations that advance large...  ...models for code. You will work with researchers to curate code examples, produce precise...  ...C and C++, Java, Rust, and Go for model training and benchmarking. Evaluate AI-generated... 
    Training
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Canada
    more than 2 months ago
  •  ...Snorkel started as a research project in the...  ...organizations to empower scientists, engineers,...  ...how AI is built? Apply to be the newest Snorkeler...  ...problems: scoping training data needs,...  ...environments, developing evaluation frameworks, and...  ...methodologies, post-training techniques... 
    Training
    Full time
    Local area

    Snorkel AI

    Canada
    19 days ago
  • $70 - $90 per hour

     ...Role Overview Apply advanced physics expertise to drive research and to teach, evaluate, and improve AI systems by converting real-world scientific knowledge into clear, high-quality training and evaluation content. This remote contractor role focuses on advanced physics... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    2 days ago
  • $80 - $90 per hour

     ...generation AI systems. You will evaluate biology content, craft and...  ...document your rationale so model training data and outputs reflect real...  ..., critical evaluation of research, and synthesis of complex subject...  ...specific acceptance criteria apply. Eligibility and Start... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    1 day ago
  • $20 - $40 per hour

     ...Role Overview Apply your biology expertise to create, review, and refine...  ...real-world biological content that trains and improves next-generation AI systems...  ...educational content. Author, evaluate, and edit laboratory reports, research papers, or field study narratives... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    12 hours ago
  • $80 - $100 per hour

     ...Role Overview Apply advanced computational physics expertise to...  ...physics tasks, influence model training data, and improve agent-...  ...improve AI model training and evaluation. Communicate findings and...  ...experience working in technical, research, or engineering contexts.... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    5 days ago
  • $100 per hour

     ...Role Overview Apply deep finance expertise to improve AI-driven...  ...standards. Develop, refine, and evaluate prompts related to financial...  ..., due-diligence reports, research summaries, and other...  ...and feedback that inform model training and evaluation. Compensation... 
    Training
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    Canada
    12 days ago
  • $40 - $50 per hour

     ...Role Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations for...  ..., clarity, and completeness. Assess LLM responses against defined instruction-following... 
    Training
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Canada
    1 day ago
  •  ...workforce development programs globally, having trained over 500,000 professionals across 11...  ...This role offers a unique opportunity to apply your data skills to design engaging and...  ...development best practices, they will research, design, and refine engaging learning experiences... 
    Training
    Full time
    Apprenticeship

    Correlation One

    Canada
    24 days ago
  • $55 - $100 per hour

     ...Role Overview Apply your supply chain expertise to create high...  ...quality, real-world input that trains next-generation AI systems....  ...chain workflows into clear data, evaluations, and reports that help models...  ...option provided with this posting to submit your information and... 
    Training
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    4 days ago
  • $120k - $160k

     ...engineers, ML engineers, and data scientists delivering the...  ...including feature engineering, training pipelines, model serving, A/...  ...focus on reproducibility, model evaluation, observability, and model lifecycle...  .../data engineering or applied data science, with 3+ years... 
    Training
    Full time
    Worldwide

    Xsolla

    Canada
    9 days ago
  • $80 - $160 per hour

     ...level condensed matter and quantum information knowledge to a research-grade physics benchmark focused on constrained quantum...  ...interpretation, and high-quality written analysis that will be used to train and evaluate next-generation AI systems. No prior AI experience is... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    2 days ago
  • $26.47 - $43.62 per hour

     ...change of address, mail holds, and giving out post office box keys. Job Overview...  ...government employment Paid on-the-job training is provided Career advancement potential...  ...area and nationwide, plus a bonus guide to applying for employment with all other government... 
    Hourly pay
    Full time
    Currently hiring
    Work at office

    Postal Jobs Source (a division of Labor Services)

    Canada
    5 minutes ago
  • $90k - $130k

     ...and indirect prompt injection. Responsibilities: We help train advanced generative systems by providing human feedback that improves...  ...large language models for real-world task execution. We evaluate system behavior carefully and provide detailed technical input... 
    Training
    Full time
    Work at office

    Outlier AI

    Canada
    more than 2 months ago
  •  ...evolving range of tasks, evaluating, stress-testing, and...  ...cases that engineering and research teams can act on....  ...and exemplars to build training and evaluation datasets...  ...standard. Build and apply rubrics and taxonomies:...  ...in AI data annotation, LLM evaluation, content moderation... 
    Training
    Part time
    Remote work
    Shift work

    Cohere

    Canada
    12 days ago
  • $140k - $165k

     ...designations and certification support, training and more! Total Rewards: Consider it coveredhealth...  ...high quality leads. Identify, evaluate, and onboard strategic channel partners (...  ...that genuinely make a difference. Youll apply advanced thinking, design expertise, and... 
    Training
    Full time
    For subcontractor
    Local area
    Flexible hours
    Shift work

    GD Mission Systems

    Canada
    10 days ago
  • $60 - $80 per hour

     ...Develop high-quality content, annotations, and explanations based on current molecular biology concepts, techniques, and findings. Evaluate AI-generated outputs for scientific validity, clarity, and alignment with established molecular biology standards. Provide... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    5 days ago
  • $100 per hour

     ...Role Overview Help train and refine advanced AI systems by applying deep software engineering expertise to evaluate, edit, and produce high‑quality...  ...proposals. Conduct independent research and fact-checking to...  ...through the original job posting. Applicants will be evaluated... 
    Training
    Remote job
    Hourly pay
    Contract work
    Part time
    For contractors

    SaidGig

    Canada
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!