Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Research Scientist, LLM Evaluation & Post-Training

Full-time

Innodata

Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.

This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.

The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.

What You’ll Own:

As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.

Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.

This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):

  • Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
  • Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
  • Develop and validate evaluation frameworks for LLM and multimodal systems, including:
    • benchmark/task design
    • scoring methods
    • judge/model-assisted evaluation
    • human evaluation protocols
    • robustness/stress testing
  • Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
  • Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
  • Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
  • Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
  • Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
  • Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
  • Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
  • Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
  • Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement

You’ll Thrive in This Role If You Have:

  • MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
  • 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
  • Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
  • Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
  • Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
  • Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
  • Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
  • Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
  • Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
Vacancy posted 27 days ago
Similar jobs that could be interesting for youBased on the Applied Research Scientist, LLM Evaluation & Post-Training in Canada vacancy
  • Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation... 
    Training
    Full time

    Innodata

    Canada
    15 days ago
  •  ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you...  .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client... 
    Training
    Full time

    Innodata

    Canada
    15 days ago
  •  ...workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you...  ...the datasets used to train, fine-tune,...  ...engagement), an Applied Research Scientist (shapes...  ...for evaluation and post-training. What You...  ...— rubric-grounded LLM-as-judge prompts, regression... 
    Training
    Full time
    Shift work

    Innodata

    Canada
    15 days ago
  •  ...Role We’re looking for Applied Scientists to join Wayve Labs and help...  ..., We Are a High-conviction Research Team With The Strategic Patience...  ...Define and evolve Evaluation Frameworks and Benchmarks for...  ...transformers, MoE, large-scale training) ~ Generative world... 
    Training
    Full time
    Work at office
    Work from home
    Visa sponsorship
    Relocation package
    Flexible hours

    Wayve

    Canada
    5 days ago
  •  ...providing the data, evaluation frameworks, and...  .... As an Applied Data Scientist, Health AI Evaluation...  ...datasets used to train, fine-tune, and evaluate...  ..., Applied Research Scientist, AI/ML...  ...for evaluation and post-training. What...  ...rubric-grounded LLM-as-judge prompts,... 
    Training
    Full time
    Shift work

    Innodata

    Canada
    15 days ago
  •  ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will...  .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client... 
    Training
    Full time

    Innodata

    Canada
    15 days ago
  • $70 - $90 per hour

     ...Role Overview Apply advanced physics expertise to help train and evaluate next-generation AI systems. This remote opportunity centers on advanced physics research, scientific reasoning, and clear communication of complex concepts. No prior AI experience is required,... 
    Training
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    23 days ago
  • $31.56 per hour

     ...is an exciting opportunity to apply your expertise while gaining...  ...equipment within the scope of training utilizing defined schedules...  ...department staff annual performance evaluations by providing objective...  ...tests. Participate in special research projects as requested;... 
    Training
    Full time
    Temporary work
    Relocation package
    Monday to Friday

    UPMC Presbyterian

    Canada
    12 days ago
  •  ...Role Overview Apply your biochemistry expertise...  ...contractor role you will evaluate biochemical data,...  ...biochemical datasets, research findings, and laboratory...  ...biochemistry practice for AI training purposes. Assess...  ...through this posting. Include your profile... 
    Training
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Canada
    23 days ago
  •  ...Role Overview Design and evaluate high-quality datasets and evaluations that advance large...  ...models for code. You will work with researchers to curate code examples, produce precise...  ...C and C++, Java, Rust, and Go for model training and benchmarking. Evaluate AI-generated... 
    Training
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Canada
    more than 2 months ago
  •  ...is hiring Medical Laboratory Scientists &Medical Lab Technicians to...  ...personal, financial, and more.  Apply today & start your career in...  ...of patients and in the evaluation of a patient’s response to treatment...  ...within the scope of training utilizing defined schedules and... 
    Training
    Full time
    Immediate start
    Relocation package

    UPMC Harrisburg

    Canada
    3 days ago
  •  ...Snorkel started as a research project in the...  ...organizations to empower scientists, engineers,...  ...how AI is built? Apply to be the newest Snorkeler...  ...problems: scoping training data needs,...  ...environments, developing evaluation frameworks, and...  ...methodologies, post-training techniques... 
    Training
    Full time
    Local area

    Snorkel AI

    Canada
    15 days ago
  • $70 - $90 per hour

     ...Role Overview Apply advanced physics expertise to drive research and to teach, evaluate, and improve AI systems by converting real-world scientific knowledge into clear, high-quality training and evaluation content. This remote contractor role focuses on advanced physics... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    3 days ago
  • $80 - $160 per hour

     ...reproducible methodological documentation to train and evaluate AI systems. Key Responsibilities...  ...of bacterial population dynamics. Apply the Euler-Lotka equation to assess and...  ...authorization restrictions were provided in the posting. How to Apply If you are interested... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    2 days ago
  • $80 - $160 per hour

     ...Role Overview Apply advanced condensed matter and magnetism expertise...  ...-quality domain inputs that train next-generation AI systems on...  ...BNS notation conventions. Evaluate optical and magneto-optic signatures...  ...example a PhD, or equivalent research experience in condensed matter... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    2 days ago
  • $80 - $90 per hour

     ...generation AI systems. You will evaluate biology content, craft and...  ...document your rationale so model training data and outputs reflect real...  ..., critical evaluation of research, and synthesis of complex subject...  ...specific acceptance criteria apply. Eligibility and Start... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    4 days ago
  • $80 - $100 per hour

     ...Role Overview Apply advanced computational physics expertise to...  ...physics tasks, influence model training data, and improve agent-...  ...improve AI model training and evaluation. Communicate findings and...  ...experience working in technical, research, or engineering contexts.... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    3 days ago
  • $20 - $40 per hour

     ...Role Overview Apply your biology expertise to create, review, and refine...  ...real-world biological content that trains and improves next-generation AI systems...  ...educational content. Author, evaluate, and edit laboratory reports, research papers, or field study narratives... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    2 days ago
  • $40 - $50 per hour

     ...Role Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations for...  ..., clarity, and completeness. Assess LLM responses against defined instruction-following... 
    Training
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Canada
    3 days ago
  •  ...workforce development programs globally, having trained over 500,000 professionals across 11...  ...This role offers a unique opportunity to apply your data skills to design engaging and...  ...development best practices, they will research, design, and refine engaging learning experiences... 
    Training
    Full time
    Apprenticeship

    Correlation One

    Canada
    27 days ago
  • $90k - $100k

     ...Office, with occasional travel to sites for training sessions Employment Type: Full-time...  ..., and participation metrics. Evaluate program outcomes and identify opportunities...  ...professional growth, and innovation. How to apply If you’re a motivated HR or business professional... 
    Training
    Full time
    Work at office

    ETRO Construction Ltd.

    Canada
    more than 2 months ago
  • $100 per hour

     ...Role Overview Apply deep finance expertise to improve AI-driven...  ...standards. Develop, refine, and evaluate prompts related to financial...  ..., due-diligence reports, research summaries, and other...  ...and feedback that inform model training and evaluation. Compensation... 
    Training
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    Canada
    15 days ago
  • $26.47 - $43.62 per hour

     ...change of address, mail holds, and giving out post office box keys. Job Overview...  ...government employment Paid on-the-job training is provided Career advancement potential...  ...area and nationwide, plus a bonus guide to applying for employment with all other government... 
    Hourly pay
    Full time
    Currently hiring
    Work at office

    Postal Jobs Source (a division of Labor Services)

    Canada
    3 days ago
  • $120k - $160k

     ...engineers, ML engineers, and data scientists delivering the...  ...including feature engineering, training pipelines, model serving, A/...  ...focus on reproducibility, model evaluation, observability, and model lifecycle...  .../data engineering or applied data science, with 3+ years... 
    Training
    Full time
    Worldwide

    Xsolla

    Canada
    12 days ago
  • $90k - $130k

     ...and indirect prompt injection. Responsibilities: We help train advanced generative systems by providing human feedback that improves...  ...large language models for real-world task execution. We evaluate system behavior carefully and provide detailed technical input... 
    Training
    Full time
    Work at office

    Outlier AI

    Canada
    more than 2 months ago
  • $60 - $80 per hour

     ...Develop high-quality content, annotations, and explanations based on current molecular biology concepts, techniques, and findings. Evaluate AI-generated outputs for scientific validity, clarity, and alignment with established molecular biology standards. Provide... 
    Training
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    8 days ago
  •  ...evolving range of tasks, evaluating, stress-testing, and...  ...cases that engineering and research teams can act on....  ...and exemplars to build training and evaluation datasets...  ...standard. Build and apply rubrics and taxonomies:...  ...in AI data annotation, LLM evaluation, content moderation... 
    Training
    Part time
    Remote work
    Shift work

    Cohere

    Canada
    15 days ago
  • $140k - $165k

     ...designations and certification support, training and more! Total Rewards: Consider it coveredhealth...  ...high quality leads. Identify, evaluate, and onboard strategic channel partners (...  ...that genuinely make a difference. Youll apply advanced thinking, design expertise, and... 
    Training
    Full time
    For subcontractor
    Local area
    Flexible hours
    Shift work

    GD Mission Systems

    Canada
    13 days ago
  •  ...Faculty Political Science Department of Applied Sciences and Professional Studies...  ...required to submit a translation/degree evaluation from a NACES approved vendor.   Who We...  ...and coursework, please visit: Faculty Training at UMGC: We are committed to your professional... 
    Training
    Full time
    Part time
    Afternoon shift

    University of Maryland Global Campus

    Canada
    20 days ago
  •  ...sales control and follow-up systems Attends product and sales training courses as requested by sales manager Keeps up-to-date on...  ...matched pension plan A fun working environment  Opportunity to grow Don't miss out on this great opportunity! Apply online today!... 
    Training
    Base plus commission
    Full time
    Local area
    Night shift

    St. Paul Dodge

    Canada
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!