Applied Research Scientist, LLM Evaluation & Post-Training
Innodata
Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.
This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.
The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.
What You’ll Own:
As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.
Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):
- Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
- Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
- Develop and validate evaluation frameworks for LLM and multimodal systems, including:
- benchmark/task design
- scoring methods
- judge/model-assisted evaluation
- human evaluation protocols
- robustness/stress testing
- Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
- Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
- Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
- Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
- Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
- Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
- Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
- Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
- Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement
You’ll Thrive in This Role If You Have:
- MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
- 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
- Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
- Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
- Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
- Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
- Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
- Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
- Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
- ...Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how...TrainingFull time
- ...Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how...TrainingFull time
- Research Scientist, LLM Evaluation & Post-Training page is loaded## Research Scientist, LLM Evaluation & Post-Traininglocations: Remote Work( USA)time type... ...collaborative research role that sits at the intersection of applied ML research, enterprise AI product development, and...TrainingFull timeRemote work
- Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross...SuggestedRemote jobHourly payFlexible hours
- ...business problems.We’re training and deploying... ...is a team of researchers, engineers, designers... ...us!Why this role?Evaluation is critical to making... ...to measure LLM progress.As a Senior Research Scientist, Model Evaluation,... ...encourage you to apply. We strive to create...TrainingFull timeWork at officeLocal areaRemote workHome office
$100 - $120 per hour
...empirical machine learning research across computer vision and... .... This role focuses on training, improving, evaluating, and deploying deep learning... ...measures such as FID; and applying training-efficiency techniques... ...-parameter generators. LLM post-training and behavioral...TrainingHourly payRemote workFlexible hours$136k
...models, large language models (LLM), generative audio (music... ...part of a close-knit team of applied scientists and product managers who are... ..., and bring cutting edge research to raise the bar within the... ...Experience with LLM model training and tuning PREFERRED QUALIFICATIONS...Training$50 per hour
...Chemistry specialists apply their subject-matter knowledge to create and evaluate prompts, review model outputs... .... Evaluate and rate LLM responses for accuracy,... ...chemistry contexts. Research chemistry topics to... ...with valid work or training authorization, for example...TrainingPart timeH1bRemote workVisa sponsorship10 hours per weekFlexible hours$50 per hour
...Overview Physics specialists apply advanced physics knowledge to evaluate and improve large... ...outputs, and conducting research-driven analyses. This part... ...role focuses on improving LLM performance across physics... ...graduates with valid work or training authorization, for...TrainingPart timeH1bRemote work10 hours per week$192k - $304.75k
...looking for a passionate scientist at the intersection... .... As a Sr. Quantum Applied Research Scientist, you will help... ...synthesis pipelines, post-trainable model... ...will span synthetic training data generation, surrogate... ..., fine-tuning, and evaluation.Strong background in...TrainingFull timeRemote work- ...you?As a Principal Applied Scientist, you will lead the architecture, research, and... ...systems, covering LLM fine-tuning, multimodal... ...services for model training, fine-tuning, large... ...and metrics for evaluating agentic systems,... ...at the top of the posted range based on the...TrainingWork at officeImmediate startRemote work
$183.8k - $248.7k
As a Senior Applied Scientist specializing in lead scoring and... ...teams to turn novel research into scalable,... ...preprocessing, distributed training, model optimization,... ...Define offline and online evaluation frameworks; establish... ...Techniques**: LLM fine-tuning, federated...TrainingFlexible hours$142.8k - $193.2k
...talented, and inventive Applied Scientist to help build industry-leading... ...understanding, modern LLM architectures, LLM evaluation & tooling, and a passion... ...assistants. * Fine-tune/post-train LLMs using techniques... ...of solutions from research to production, including...TrainingFlexible hours$119.8k - $234.7k
...ContributorTravel: Less than 25%Profession: Research, Applied, & Data SciencesDiscipline:... ..., including the latest LLM models. Deploying robust and... ...Experience with distributed training or inference for SLMs and... ...implementation. Experience evaluating SLMs, LLMs, or agentic...TrainingOngoing contractWork experience placementWork at officeLocal area$126.2k - $264.1k
...As our Principal Applied Scientist, you will play a key... ...drive projects from research POC to production.... ...including data, model, training, and evaluation, employing best practices... ...technologies in LLM and generative AI, such... ...provided in this posting are specific to the...TrainingTemporary workRemote workFlexible hours- ...Applied Scientist, Alexa International Tech Alexa International... ...to scientific research and applied AI for multi... ...language processing, modern LLM architectures, LLM evaluation & tooling, and a... ...-Speech (S2S) model training and fine-tuning Fine-tune/post-train LLMs using techniques...Training
- ...autonomously. As a Principal Applied Scientist, you'll help invent... ...You will lead the research and development of... ...workflows, and post-training techniques to build autonomous... ..., training and evaluating large models, developing... ...model training LLM post-training Reinforcement...TrainingFull timeWork at officeImmediate startRemote work
- ...skilled and driven Senior Applied Scientist with 7+ years of industry experience... .... Responsibilities ~Research and engineer NLP solutions... ...and implement algorithms, train state of the art large language models (LLM) on large data, and evaluate their performance. ~...TrainingFull time
- ...estimator does. As a Computer Vision Applied Research Scientist at Boon, you will own end-to-end... ...Research & Architecture ~Design and evaluate novel multi-stage vision architectures... ..., fusion strategies, loss functions, training regimes. ~Run rigorous experiments...TrainingFull time
$147.6k - $274.2k
...sits within the applied science... ...~8+ years of post-degree industry... ...production — not research-only experience... ...unstructured text ~LLM-based information... ..., and post-training ~Knowledge distillation... ...~End-to-end evaluation framework... ...Mentor applied scientists and ML practitioners...TrainingFull timeFlexible hours$100 - $120 per hour
...LLM Research Scientist (Pre-training & Computer Vision & Adversarial Robustness) is a remote red-team track... ...Why this role matters Adversarial evaluation is how AuraOne hardens AI models before... ...$120 / hr Application process Apply through AuraOne's specialist intake...TrainingRemote jobFor contractors10 hours per week$100 - $120 per hour
...-scoped, empirical ML research projects that advance... ...model development, from training models from scratch to... ...and TRADES. Evaluating robust accuracy under... ...metrics like FID. Applying training-efficiency techniques... ...parameter counts. LLM Post-Training and Behavioral...TrainingHourly payTemporary workRemote workFlexible hours$100 - $120 per hour
...LLM Research Scientist (Pre-training & Post-Training) is a remote review track for evaluating AI outputs across research science reasoning, calculations, and research workflows.... ...Graduate-level training or equivalent applied experience in research science or a closely...TrainingRemote jobFor contractors10 hours per week- ...Overview Use your Biology training to design and solve... ..., help define evaluation benchmarks across undergraduate... ...with model researchers to probe and improve model... ...explanations. Work with LLM researchers to align... ...eligible and encouraged to apply. Applicants must...TrainingContract workFor contractorsFreelanceRemote work
$193.93k - $352.29k
...output can be trusted — evaluation, verification, and the... ...standard of proof we apply to the vehicle. Our... ...of every engineer and researcher at Nuro by 100x. Not a... ...that means nothing. Post-train models on data nobody... ...Direct experience with LLM agent systems — building...TrainingFull time$167.1k - $226.1k
...do. You will distill LLM quality judgments into... ...into online objectives. Scientists on this team own the... ...combination of problems: research-grade modeling with an... ...online. That means training compact models against... ...degree and 6+ years of applied research experience- Experience...TrainingFlexible hours$137.7k - $275.4k
...Incubation team seeks a Research Scientist to advance AI-driven... ...group of PhDs and applied scientists working at... ...operating at the frontier of LLM posttraining, agentic... ...architectures and post-training strategies that... ...contexts.Developing evaluation and benchmarking frameworks...TrainingFull timeWork at officeRemote workWorldwide$142.8k - $193.2k
...your needs.Key job responsibilitiesAs an Applied Scientist on the team, you will help lead science... ...developing state of the art thinking-LLM-based techniques to reason about customers... ...cross-attentive LLM rankers ; training multi-objective ranking and optimization...TrainingFlexible hours$119.8k - $234.7k
...than 25%Profession: Research, Applied, & Data SciencesDiscipline... ...through clicks, post-click engagement, and... ...outcomes. We design and train transformer-based... ...rigorous offline and online evaluation. We build... ...dynamics.Engineers and scientists on the team work at the...TrainingOngoing contractWork at officeLocal areaShift work$117.13k - $135k
Research Scientist Center for Effective OrganizationUSC... ...from other applied research centers... ...assessment, and evaluating the effectiveness... ...skillsProficiency in use of LLM toolsExpertise in... ..., education/training, key skills, internal... ...Job openings are posted for a minimum of...TrainingFull timeWork experience placementLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!







