Applied Research Scientist, LLM Evaluation & Post-Training
Innodata
Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.
This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.
The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.
What You’ll Own:
As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.
Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):
- Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
- Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
- Develop and validate evaluation frameworks for LLM and multimodal systems, including:
- benchmark/task design
- scoring methods
- judge/model-assisted evaluation
- human evaluation protocols
- robustness/stress testing
- Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
- Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
- Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
- Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
- Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
- Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
- Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
- Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
- Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement
You’ll Thrive in This Role If You Have:
- MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
- 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
- Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
- Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
- Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
- Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
- Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
- Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
- Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
- Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation...TrainingFull time
- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you... .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client...TrainingFull time
- ...workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you... ...the datasets used to train, fine-tune,... ...engagement), an Applied Research Scientist (shapes... ...for evaluation and post-training. What You... ...— rubric-grounded LLM-as-judge prompts, regression...TrainingFull timeShift work
- ...Role We’re looking for Applied Scientists to join Wayve Labs and help... ..., We Are a High-conviction Research Team With The Strategic Patience... ...Define and evolve Evaluation Frameworks and Benchmarks for... ...transformers, MoE, large-scale training) ~ Generative world...TrainingFull timeWork at officeWork from homeVisa sponsorshipRelocation packageFlexible hours
- ...providing the data, evaluation frameworks, and... .... As an Applied Data Scientist, Health AI Evaluation... ...datasets used to train, fine-tune, and evaluate... ..., Applied Research Scientist, AI/ML... ...for evaluation and post-training. What... ...rubric-grounded LLM-as-judge prompts,...TrainingFull timeShift work
- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will... .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client...TrainingFull time
$70 - $90 per hour
...Role Overview Apply advanced physics expertise to help train and evaluate next-generation AI systems. This remote opportunity centers on advanced physics research, scientific reasoning, and clear communication of complex concepts. No prior AI experience is required,...TrainingHourly payFor contractorsRemote work$31.56 per hour
...is an exciting opportunity to apply your expertise while gaining... ...equipment within the scope of training utilizing defined schedules... ...department staff annual performance evaluations by providing objective... ...tests. Participate in special research projects as requested;...TrainingFull timeTemporary workRelocation packageMonday to Friday- ...Role Overview Apply your biochemistry expertise... ...contractor role you will evaluate biochemical data,... ...biochemical datasets, research findings, and laboratory... ...biochemistry practice for AI training purposes. Assess... ...through this posting. Include your profile...TrainingHourly payFor contractorsRemote work
- ...Role Overview Design and evaluate high-quality datasets and evaluations that advance large... ...models for code. You will work with researchers to curate code examples, produce precise... ...C and C++, Java, Rust, and Go for model training and benchmarking. Evaluate AI-generated...TrainingFor contractorsRemote work10 hours per weekFlexible hours
- ...is hiring Medical Laboratory Scientists &Medical Lab Technicians to... ...personal, financial, and more. Apply today & start your career in... ...of patients and in the evaluation of a patient’s response to treatment... ...within the scope of training utilizing defined schedules and...TrainingFull timeImmediate startRelocation package
- ...Snorkel started as a research project in the... ...organizations to empower scientists, engineers,... ...how AI is built? Apply to be the newest Snorkeler... ...problems: scoping training data needs,... ...environments, developing evaluation frameworks, and... ...methodologies, post-training techniques...TrainingFull timeLocal area
$70 - $90 per hour
...Role Overview Apply advanced physics expertise to drive research and to teach, evaluate, and improve AI systems by converting real-world scientific knowledge into clear, high-quality training and evaluation content. This remote contractor role focuses on advanced physics...TrainingRemote jobHourly payFor contractors$80 - $160 per hour
...reproducible methodological documentation to train and evaluate AI systems. Key Responsibilities... ...of bacterial population dynamics. Apply the Euler-Lotka equation to assess and... ...authorization restrictions were provided in the posting. How to Apply If you are interested...TrainingRemote jobHourly payFor contractors$80 - $160 per hour
...Role Overview Apply advanced condensed matter and magnetism expertise... ...-quality domain inputs that train next-generation AI systems on... ...BNS notation conventions. Evaluate optical and magneto-optic signatures... ...example a PhD, or equivalent research experience in condensed matter...TrainingRemote jobHourly payFor contractors$80 - $90 per hour
...generation AI systems. You will evaluate biology content, craft and... ...document your rationale so model training data and outputs reflect real... ..., critical evaluation of research, and synthesis of complex subject... ...specific acceptance criteria apply. Eligibility and Start...TrainingRemote jobHourly payFor contractors$80 - $100 per hour
...Role Overview Apply advanced computational physics expertise to... ...physics tasks, influence model training data, and improve agent-... ...improve AI model training and evaluation. Communicate findings and... ...experience working in technical, research, or engineering contexts....TrainingRemote jobHourly payFor contractors$20 - $40 per hour
...Role Overview Apply your biology expertise to create, review, and refine... ...real-world biological content that trains and improves next-generation AI systems... ...educational content. Author, evaluate, and edit laboratory reports, research papers, or field study narratives...TrainingRemote jobHourly payFor contractors$40 - $50 per hour
...Role Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations for... ..., clarity, and completeness. Assess LLM responses against defined instruction-following...TrainingRemote jobHourly payFor contractorsImmediate start- ...workforce development programs globally, having trained over 500,000 professionals across 11... ...This role offers a unique opportunity to apply your data skills to design engaging and... ...development best practices, they will research, design, and refine engaging learning experiences...TrainingFull timeApprenticeship
$90k - $100k
...Office, with occasional travel to sites for training sessions Employment Type: Full-time... ..., and participation metrics. Evaluate program outcomes and identify opportunities... ...professional growth, and innovation. How to apply If you’re a motivated HR or business professional...TrainingFull timeWork at office$100 per hour
...Role Overview Apply deep finance expertise to improve AI-driven... ...standards. Develop, refine, and evaluate prompts related to financial... ..., due-diligence reports, research summaries, and other... ...and feedback that inform model training and evaluation. Compensation...TrainingHourly payPart timeFor contractorsRemote work$26.47 - $43.62 per hour
...change of address, mail holds, and giving out post office box keys. Job Overview... ...government employment Paid on-the-job training is provided Career advancement potential... ...area and nationwide, plus a bonus guide to applying for employment with all other government...Hourly payFull timeCurrently hiringWork at office$120k - $160k
...engineers, ML engineers, and data scientists delivering the... ...including feature engineering, training pipelines, model serving, A/... ...focus on reproducibility, model evaluation, observability, and model lifecycle... .../data engineering or applied data science, with 3+ years...TrainingFull timeWorldwide$90k - $130k
...and indirect prompt injection. Responsibilities: We help train advanced generative systems by providing human feedback that improves... ...large language models for real-world task execution. We evaluate system behavior carefully and provide detailed technical input...TrainingFull timeWork at office$60 - $80 per hour
...Develop high-quality content, annotations, and explanations based on current molecular biology concepts, techniques, and findings. Evaluate AI-generated outputs for scientific validity, clarity, and alignment with established molecular biology standards. Provide...TrainingRemote jobHourly payFor contractors- ...evolving range of tasks, evaluating, stress-testing, and... ...cases that engineering and research teams can act on.... ...and exemplars to build training and evaluation datasets... ...standard. Build and apply rubrics and taxonomies:... ...in AI data annotation, LLM evaluation, content moderation...TrainingPart timeRemote workShift work
$140k - $165k
...designations and certification support, training and more! Total Rewards: Consider it coveredhealth... ...high quality leads. Identify, evaluate, and onboard strategic channel partners (... ...that genuinely make a difference. Youll apply advanced thinking, design expertise, and...TrainingFull timeFor subcontractorLocal areaFlexible hoursShift work- ...Faculty Political Science Department of Applied Sciences and Professional Studies... ...required to submit a translation/degree evaluation from a NACES approved vendor. Who We... ...and coursework, please visit: Faculty Training at UMGC: We are committed to your professional...TrainingFull timePart timeAfternoon shift
- ...sales control and follow-up systems Attends product and sales training courses as requested by sales manager Keeps up-to-date on... ...matched pension plan A fun working environment Opportunity to grow Don't miss out on this great opportunity! Apply online today!...TrainingBase plus commissionFull timeLocal areaNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!



