Applied Research Scientist, LLM Evaluation & Post-Training
Innodata
Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.
This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.
The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.
What You’ll Own:
As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.
Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):
- Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
- Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
- Develop and validate evaluation frameworks for LLM and multimodal systems, including:
- benchmark/task design
- scoring methods
- judge/model-assisted evaluation
- human evaluation protocols
- robustness/stress testing
- Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
- Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
- Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
- Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
- Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
- Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
- Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
- Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
- Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement
You’ll Thrive in This Role If You Have:
- MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
- 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
- Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
- Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
- Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
- Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
- Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
- Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
- Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
- Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation...TrainingFull time
- ...Role We’re looking for Applied Scientists to join Wayve Labs and help... ..., We Are a High-conviction Research Team With The Strategic Patience... ...Define and evolve Evaluation Frameworks and Benchmarks for... ...transformers, MoE, large-scale training) ~ Generative world...TrainingFull timeWork at officeWork from homeVisa sponsorshipRelocation packageFlexible hours
- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will... .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client...TrainingFull time
- ...Snorkel started as a research project in the... ...organizations to empower scientists, engineers,... ...how AI is built? Apply to be the newest Snorkeler... ...problems: scoping training data needs,... ...environments, developing evaluation frameworks, and... ...methodologies, post-training techniques...TrainingFull timeLocal area
- ...Role Overview Design and evaluate high-quality datasets and evaluations that advance large... ...models for code. You will work with researchers to curate code examples, produce precise... ...C and C++, Java, Rust, and Go for model training and benchmarking. Evaluate AI-generated...TrainingFor contractorsRemote work10 hours per weekFlexible hours
- ...Position Description The Quality Training Coordinator will own the day-to... ...new and revised procedures are evaluated for training impacts; and... ...of ANSI/ANS-15.8 and experience applying its quality assurance expectations in a research reactor, test reactor, advanced...TrainingFull timeTemporary workFor contractorsTraineeshipRemote work
$26.47 - $43.62 per hour
...change of address, mail holds, and giving out post office box keys. Job Overview... ...government employment Paid on-the-job training is provided Career advancement potential... ...area and nationwide, plus a bonus guide to applying for employment with all other government...Hourly payFull timeCurrently hiringWork at office$90k - $130k
...and indirect prompt injection. Responsibilities: We help train advanced generative systems by providing human feedback on AI agents... ...and provide detailed technical observations. We evaluate agent interactions and identify failures that may affect safety...TrainingFull time- ...evolving range of tasks, evaluating, stress-testing, and... ...cases that engineering and research teams can act on.... ...and exemplars to build training and evaluation datasets... ...standard. Build and apply rubrics and taxonomies:... ...in AI data annotation, LLM evaluation, content moderation...TrainingPart timeRemote workShift work
- ...the Role This is a dedicated growth research role. You'll run studies — usability testing... ...studies and instruments to conduct evaluation and diagnostic discovery for adoption experiences... ...experience in UX research or related applied research field ~ Strong quantitative...Full timeImmediate start
- ...Faculty Political Science Department of Applied Sciences and Professional Studies... ...required to submit a translation/degree evaluation from a NACES approved vendor. Who We... ...and coursework, please visit: Faculty Training at UMGC: We are committed to your professional...TrainingFull timePart timeAfternoon shift
$52 - $56 per hour
...Senior Analytical Scientist - ICP-MS JOB-10047520 Anticipated... ...safety as their priority, training will be provided for all... ...direction Perform analysis and research with minimal supervision... ...will be offered within this posted range based on experience, skills...TrainingFull timeContract workTemporary workLocal areaImmediate start- ...Generous paid time off, company paid training and tuition reimbursement. ~ Positive and... ...To learn more about our company, and to apply online for this exciting opportunity, visit... ...65279 Category: Plant & Facilities Posting Date: 2026-08-25 Job Schedule: Full time...Permanent employmentFull time
$45 per hour
...Generous paid time off, company paid training and tuition reimbursement. Positive and... ...To learn more about our company, and to apply online for this exciting opportunity, visit... ...65276 Category: Plant & Facilities Posting Date: 2026-08-27 Job Schedule: Full time...Permanent employmentFull timeApprenticeship$36.02 per hour
...JOB POSTING Project Lead, Choice for Change (Part-time Contract until March 31, 202... ...before a charge is laid. If you want to apply your skills, experience, and talents to... ...from recognized institution or equivalent training in community work sectors Minimum of 2...TrainingHourly payContract workPart timeWork experience placementAfternoon shift$38 - $45 per hour
...Generous paid time off, company paid training and tuition reimbursement. Positive and... ...To learn more about our company, and to apply online for this exciting opportunity, visit... ...65424 Category: Plant & Facilities Posting Date: 2026-09-01 Job Schedule: Full time...Full timeApprenticeshipCasual workMonday to Friday$70k - $150k
...engineering and techniques for evaluating output quality. We need... ...production experience supporting an LLM-based feature used by end... ...regressions before release. You will apply safety and security controls... .... We support our people with training, coaching, manager support,...TrainingFull time$24 per hour
...energy. With safety as their priority, training will be provided for all employees to ensure... ...a willingness to grow are encouraged to apply. Job Description Maintain and... ...transfer procedures Provide initial evaluations and obtain assistance for technical problem...TrainingFull timeContract workTemporary workFor contractorsLocal areaMonday to FridayNight shiftDay shift$298k
...systems including LLM-powered copilots,... ...model development to evaluation, deployment, and... ...implementation of model training and fine-tuning... ...ML engineers, ML scientists, and AI architects... ...at the time of posting and may be updated... ...Eligibility requirements apply to some benefits....TrainingFull timeWork at office- ...Applicants with the drive to succeed, a strong work ethic and always seek to improve are encouraged to apply. Pipeline (or similar) experience is preferred but additional training will be provided for suitable candidates new to the energy infrastructure industry. KEY...TrainingDaily paidPermanent employmentFull timeWork at officeShift workDay shift
- ...platforms, with a strong focus on LLM applications, agentic AI systems, applied machine learning, backend engineering, data pipelines, evaluation, and production operations. You will... ...~5+ years of end-to-end experience training, evaluating, testing, deploying, and...TrainingFull timeFlexible hours
$50 - $70 per hour
...Help improve frontier AI systems by evaluating the everyday professional materials they produce. You will assess documents, presentations... ...This role is designed for detail-oriented generalists who can apply sound judgment across varied subjects and formats. Deep specialization...TrainingHourly payRemote work$41.68 - $43.98 per hour
...JOB POSTING Manager, SWIS Program (Full-time, Permanent)... ...communities. If you want to apply your skills, experience, and... ...overall delivery, monitoring and evaluation of the Settlement Workers in... ...counselling diploma or equivalent training in Human Services field...TrainingPermanent employmentFull timeTemporary workSummer workWork at officeLocal areaAfternoon shift$330k
...program , including regular ‘step change’ training and frequent AI demos relevant to each... ...AR$ 141,050,000 in Argentina . Our evaluation process is as follows: Interview with... ...in this role, we encourage you to apply. Change.org is an open platform designed...TrainingFull time- ...connect with guests to ensure exceptional satisfaction, implement sales initiatives, oversee front-of-house, and assist in hiring and training. Join our team as an Assistant General Manager and immerse yourself in a role where you will learn and grow every day. We offer...TrainingFull timeImmediate start
$60 per hour
...Benefits: - Competitive wages up to $60/hour based on specialty training (we provide training if you don't have it!). - Housing... ...ready to embrace a rewarding career and an enviable lifestyle, apply today and take the first step toward joining our dynamic team!...TrainingHourly payPrice workFull timeApprenticeshipCasual workSeasonal workLive inLocal areaRelocation packageMonday to FridayFlexible hours$30 - $40 per hour
...schedules into the overall project schedule and apply progress updates. Understand the... ...project life cycle. Assist with the evaluation of actual construction progress,... ...or a relevant combination of education / training and field experience. Experience...TrainingHourly payTemporary workWork experience placementRemote workShift workDay shift- ...experiment to production — feature stores, training pipelines, model serving, and monitoring.... ...) as a force multiplier. And you build LLM-powered pipelines and autonomous agents that... ...volume you designed for. ~ Experience applying LLMs and agentic systems in production...TrainingFull time
$40.5 per hour
...with support from management, a tool allowance, and manufacturer training! Once you're enrolled in the apprentice program we'll even pay... ...Allowance signing bonus, build your box on us! Interested? Apply online today or come in and see Brad at Bonnyville Dodge!...TrainingFull timeApprenticeshipImmediate startRelocation package$68k
...Two postdoctoral positions are available for highly motivated researchers interested in human immunology, aging and immune mechanisms that... ..., or computational analysis are especially encouraged to apply. Application materials should include: Cover letter describing...Full timeFixed term contract
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!





