Applied Research Scientist, LLM Evaluation & Post-Training
Innodata
Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.
This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.
The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.
What You’ll Own:
As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.
Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):
- Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
- Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
- Develop and validate evaluation frameworks for LLM and multimodal systems, including:
- benchmark/task design
- scoring methods
- judge/model-assisted evaluation
- human evaluation protocols
- robustness/stress testing
- Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
- Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
- Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
- Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
- Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
- Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
- Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
- Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
- Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement
You’ll Thrive in This Role If You Have:
- MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
- 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
- Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
- Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
- Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
- Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
- Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
- Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
- Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
- Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation...TrainingFull time
- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you... .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client...TrainingFull time
- ...Role We’re looking for Applied Scientists to join Wayve Labs and help... ..., We Are a High-conviction Research Team With The Strategic Patience... ...Define and evolve Evaluation Frameworks and Benchmarks for... ...transformers, MoE, large-scale training) ~ Generative world...TrainingFull timeWork at officeWork from homeVisa sponsorshipRelocation packageFlexible hours
- ...providing the data, evaluation frameworks, and... .... As an Applied Data Scientist, Health AI Evaluation... ...datasets used to train, fine-tune, and evaluate... ..., Applied Research Scientist, AI/ML... ...for evaluation and post-training. What... ...rubric-grounded LLM-as-judge prompts,...TrainingFull timeShift work
- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will... .... You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client...TrainingFull time
$80 - $90 per hour
...Role Overview Apply advanced biology expertise to help train and evaluate next-generation AI systems. You will provide scientifically rigorous, real-world input... ...Experience analyzing scientific data, evaluating research critically, and synthesizing complex material....TrainingHourly payFor contractorsRemote work- ...workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you... ...the datasets used to train, fine-tune,... ...engagement), an Applied Research Scientist (shapes... ...for evaluation and post-training. What You... ...— rubric-grounded LLM-as-judge prompts, regression...TrainingFull timeShift work
$80 - $100 per hour
...Role Overview Apply deep computational physics knowledge to produce detailed, real... ...technical feedback that improve model training and evaluation. Communicate findings accurately... ...record of technical problem solving or research. Proficiency in Python and comfort...TrainingHourly payFor contractorsRemote work$80 - $160 per hour
...condensed matter and magnetism analysis to a research program focused on magnetic space... ...high quality, real-world content that trains and evaluates next generation AI models. No prior... ...knowledge of neutron scattering techniques as applied to magnetic structure determination....TrainingHourly payFor contractorsRemote work$20 - $40 per hour
...Role Overview Apply your biology expertise to improve... ...; your academic and research background is the key... ...content. Author, evaluate, and edit laboratory reports... ...materials used for AI training and validation.... ...sponsorship details are provided in this posting....TrainingHourly payContract workFor contractorsRemote workVisa sponsorship- ...Role Overview Apply your biochemistry expertise... ...contractor role you will evaluate biochemical data,... ...biochemical datasets, research findings, and laboratory... ...biochemistry practice for AI training purposes. Assess... ...through this posting. Include your profile...TrainingHourly payFor contractorsRemote work
$80 - $150 per hour
...Role Overview Apply deep physics expertise to evaluate, verify, and improve scientific... ...reasoning used to train next-generation AI... ...produced by researchers or AI systems. Detect... ...or senior research scientist. Recent... ...qualification. This posting is currently closed...Hourly payFor contractorsRemote work- ...Role Overview Design and evaluate high-quality datasets and evaluations that advance large... ...models for code. You will work with researchers to curate code examples, produce precise... ...C and C++, Java, Rust, and Go for model training and benchmarking. Evaluate AI-generated...TrainingFor contractorsRemote work10 hours per weekFlexible hours
- ...Snorkel started as a research project in the... ...organizations to empower scientists, engineers,... ...how AI is built? Apply to be the newest Snorkeler... ...problems: scoping training data needs,... ...environments, developing evaluation frameworks, and... ...methodologies, post-training techniques...TrainingFull timeLocal area
$70 - $90 per hour
...Role Overview Apply advanced physics expertise to drive research and to teach, evaluate, and improve AI systems by converting real-world scientific knowledge into clear, high-quality training and evaluation content. This remote contractor role focuses on advanced physics...TrainingRemote jobHourly payFor contractors$80 - $90 per hour
...generation AI systems. You will evaluate biology content, craft and... ...document your rationale so model training data and outputs reflect real... ..., critical evaluation of research, and synthesis of complex subject... ...specific acceptance criteria apply. Eligibility and Start...TrainingRemote jobHourly payFor contractors$20 - $40 per hour
...Role Overview Apply your biology expertise to create, review, and refine... ...real-world biological content that trains and improves next-generation AI systems... ...educational content. Author, evaluate, and edit laboratory reports, research papers, or field study narratives...TrainingRemote jobHourly payFor contractors$80 - $100 per hour
...Role Overview Apply advanced computational physics expertise to... ...physics tasks, influence model training data, and improve agent-... ...improve AI model training and evaluation. Communicate findings and... ...experience working in technical, research, or engineering contexts....TrainingRemote jobHourly payFor contractors$100 per hour
...Role Overview Apply deep finance expertise to improve AI-driven... ...standards. Develop, refine, and evaluate prompts related to financial... ..., due-diligence reports, research summaries, and other... ...and feedback that inform model training and evaluation. Compensation...TrainingHourly payPart timeFor contractorsRemote work$40 - $50 per hour
...Role Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations for... ..., clarity, and completeness. Assess LLM responses against defined instruction-following...TrainingRemote jobHourly payFor contractorsImmediate start- ...workforce development programs globally, having trained over 500,000 professionals across 11... ...This role offers a unique opportunity to apply your data skills to design engaging and... ...development best practices, they will research, design, and refine engaging learning experiences...TrainingFull timeApprenticeship
$55 - $100 per hour
...Role Overview Apply your supply chain expertise to create high... ...quality, real-world input that trains next-generation AI systems.... ...chain workflows into clear data, evaluations, and reports that help models... ...option provided with this posting to submit your information and...TrainingHourly payFor contractorsRemote work$120k - $160k
...engineers, ML engineers, and data scientists delivering the... ...including feature engineering, training pipelines, model serving, A/... ...focus on reproducibility, model evaluation, observability, and model lifecycle... .../data engineering or applied data science, with 3+ years...TrainingFull timeWorldwide$80 - $160 per hour
...level condensed matter and quantum information knowledge to a research-grade physics benchmark focused on constrained quantum... ...interpretation, and high-quality written analysis that will be used to train and evaluate next-generation AI systems. No prior AI experience is...TrainingRemote jobHourly payFor contractors$26.47 - $43.62 per hour
...change of address, mail holds, and giving out post office box keys. Job Overview... ...government employment Paid on-the-job training is provided Career advancement potential... ...area and nationwide, plus a bonus guide to applying for employment with all other government...Hourly payFull timeCurrently hiringWork at office$90k - $130k
...and indirect prompt injection. Responsibilities: We help train advanced generative systems by providing human feedback that improves... ...large language models for real-world task execution. We evaluate system behavior carefully and provide detailed technical input...TrainingFull timeWork at office- ...evolving range of tasks, evaluating, stress-testing, and... ...cases that engineering and research teams can act on.... ...and exemplars to build training and evaluation datasets... ...standard. Build and apply rubrics and taxonomies:... ...in AI data annotation, LLM evaluation, content moderation...TrainingPart timeRemote workShift work
$140k - $165k
...designations and certification support, training and more! Total Rewards: Consider it coveredhealth... ...high quality leads. Identify, evaluate, and onboard strategic channel partners (... ...that genuinely make a difference. Youll apply advanced thinking, design expertise, and...TrainingFull timeFor subcontractorLocal areaFlexible hoursShift work$60 - $80 per hour
...Develop high-quality content, annotations, and explanations based on current molecular biology concepts, techniques, and findings. Evaluate AI-generated outputs for scientific validity, clarity, and alignment with established molecular biology standards. Provide...TrainingRemote jobHourly payFor contractors$100 per hour
...Role Overview Help train and refine advanced AI systems by applying deep software engineering expertise to evaluate, edit, and produce high‑quality... ...proposals. Conduct independent research and fact-checking to... ...through the original job posting. Applicants will be evaluated...TrainingRemote jobHourly payContract workPart timeFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!



