Applied Research Scientist, LLM Evaluation & Post-Training
Innodata
Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.
This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.
The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.
What You’ll Own:
As an Applied Research Scientist, LLM Evaluation & Post-Training , you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.
Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):
- Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
- Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
- Develop and validate evaluation frameworks for LLM and multimodal systems, including:
- benchmark/task design
- scoring methods
- judge/model-assisted evaluation
- human evaluation protocols
- robustness/stress testing
- Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
- Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
- Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
- Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
- Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
- Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
- Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
- Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
- Contribute to thought leadership and best practices in LLM evaluation , post-training, and GenAI quality measurement
You’ll Thrive in This Role If You Have:
- MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field ( PhD strongly preferred )
- 5+ years of relevant experience in applied research / research science in ML/AI , with substantial work in LLMs or foundation models
- Demonstrated experience with LLM evaluation , benchmarking, alignment, post-training, or model quality research
- Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
- Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
- Experience working with modern ML tooling/frameworks (e.g., PyTorch , Hugging Face , JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
- Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
- Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
- Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
$175k - $225k
Role Description Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation...TrainingFull time$126.2k - $264.1k
...As our Principal Applied Scientist, you will play a key... ...drive projects from research POC to production.... ...including data, model, training, and evaluation, employing best practices... ...technologies in LLM and generative AI, such... ...provided in this posting are specific to the...TrainingTemporary workRemote workFlexible hours$100 - $120 per hour
...LLM Research Scientist (Pre-training & Post-Training) is a remote review track for evaluating AI outputs across research science reasoning, calculations, and research workflows.... ...Graduate-level training or equivalent applied experience in research science or a closely...TrainingRemote jobFor contractors10 hours per week- ...power grids, buildings, trains, hospitals. Industrial... ...major opportunity in applied AI, and one of the... ...intersection of machine learning research, real world data, and... .... As Applied Scientist, you work on the science... ...control, planning, or evaluation. Take problems from ambiguous...TrainingLocal area
- ...The Role We're looking for an Applied Scientist who thrives at the intersection of applied research and real-world products. You'll... ...in data, interaction, and evaluation with both creativity and engineering... ...of data modeling for training and how curation decisions shape...TrainingFlexible hours
$167.1k - $226.1k
...and Catalog Systems (ASCS) – Applied Scientist At Amazon Selection and... ...groceries to digital content. The research challenges are immense.... ...data with rigorous evaluation frameworks Define research... ...Qualifications Experience with training and deploying machine learning...TrainingWorldwide$80 - $150 per hour
...future of AI systems by applying your physics expertise. As... ...play a critical role in evaluating and enhancing the training of next-generation AI models... ...arguments generated by researchers or AI platforms. Detect... ..., or senior research scientist. Recent (last ~5 years...TrainingRemote jobHourly payFor contractors$75 per hour
...Role Overview Physics experts apply advanced physics training to evaluate AI-generated scientific content and provide detailed feedback that improves... ...full-time or tenure-track role, suitable to combine with research, teaching, postdoctoral work, or industry employment....TrainingHourly payFull timeContract workPart timeRemote workFlexible hours$88.6k - $150.8k
Join to apply for the Research Scientist role at The Henry M. Jackson Foundation for... ...clinical tool development and training, cognitive monitoring and... ...on employee performance evaluations. Qualifications Education... ...Learning Scientist, NLP/LLM San Francisco, CA $140,000...TrainingFull timeContract workFor contractorsFixed term contractInternshipWork at officeLocal areaRemote workFlexible hours- ...skilled and driven Senior Applied Scientist with 7+ years of industry experience... .... Responsibilities ~Research and engineer NLP solutions... ...and implement algorithms, train state of the art large language models (LLM) on large data, and evaluate their performance. ~...TrainingFull time
$60 - $90 per hour
...collaborate closely with researchers to convert findings into robust evaluation benchmarks. Key... ...familiarity with LLM capabilities,... ...in AI model training, model evaluation,... ...disabilities. How to Apply If you found this... ...or the original posting. The employer-of-record...TrainingHourly payFull timeFreelanceRemote work- ...record of exceptional research or engineering achievement... ...… … you should apply for this role! About... ...learning, and small custom post-trained models (SFT and RLVR)... ...AI Research Scientist to join our small team... ...from data generation to evaluation to product integration...TrainingFull timeRelocation package
$40 per hour
...We are looking for an Applied Mathematician to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level of...Hourly payFull timeContract workPart timeRemote work$150k - $230k
...Engineer to drive the post-training of our large language... ...large GPU clusters , applying distributed-training... ...Build and maintain evaluation and reward/verifier pipelines... ...with post-training research and turn promising... ...Requirements Hands-on LLM post-training...TrainingFull timeLocal areaWork from home$75 per hour
...Role Overview AI & Machine Learning Researchers apply advanced computer science and machine learning research expertise to evaluate AI-generated responses and produce expert-level training data that improves AI understanding of contemporary CS and AI/ML research and...TrainingHourly payContract workPart timeRemote workFlexible hours- A leading data services company is seeking an Applied Mathematician to evaluate AI models by providing complex mathematical problems to chatbots and assessing their outputs for quality and performance. This role offers flexibility with fully remote work and allows you to...Hourly payRemote work
- ...frontier of machine learning research. Founded by leading... ...systems, and advanced training methodologies. These... ...large-scale pre-training, post-training, reinforcement... ...and data curation. Evaluation, benchmarking, reward modeling... ...used at scale. Why Apply You’ll gain access to opportunities...TrainingCurrently hiring
$204k - $259k
...service and can also be applied to a range of vehicle... ...with other research teams in Alphabet. AI... ...learning, and robust evaluation. Role Summary In this... ...report to a Principal Scientist. Responsibilities Participate... ...World Model post‑training and evaluation Research...TrainingTemporary workRemote work$100k
...Research Scientist Shift Type: -- ; Education: Doctorate ; Location: PRE... ...national presence in the area of applied science, with funded... ...implementation of complex program evaluation strategies and the conduct... ...and scientists provide training and technical assistance in...TrainingFull timeTemporary workPart timeRemote workFlexible hoursShift work$92.15k
...1/2026; Rank CR - Research Asst Professor; Working... ...Title Research Scientist; Position Number D... ..., research, and training unit for population... ...to collect and evaluate all necessary data... ...of experience in applied social science research... ...Search Details Posting Close Date; Projected...TrainingPermanent employmentFixed term contractRemote workRelocation$175k - $215k
...hail service and can also be applied to a range of vehicle platforms... ...will: Design, implement, and evaluate state‑of‑the‑art generative... ...tested code to bring cutting‑edge research into production Partner with... ..., experience, relevant training and education, and skill level...TrainingFull timeInternshipRemote work$183.8k - $248.7k
...ads. Key Job Responsibilities Lead the research and development of ML models that personalize... ...analyze large‑scale A/B experiments, applying causal inference techniques to measure... ...model fundamentals—including architecture, training/inference lifecycles, and optimization...TrainingFlexible hours- ...is a non-profit AI research institute dedicated... ...Conducting pre- and post-release adversarial evaluations of frontier models... ...fine-tuning and post-training workflows to... ...significant overlap between scientist and engineer roles.... ...is by applying directly via the application...TrainingFull timeRemote workVisa sponsorship
$183.8k - $248.7k
...the full ML lifecycle, from exploratory research and offline modeling to online experimentation... ...analyze large-scale A/B experiments, applying causal inference techniques to measure... ...fundamentals, including architecture, training/inference lifecycles, and optimization of...TrainingFlexible hoursShift work- ...The Role We’re looking for Applied Scientists to join Wayve Labs and help... ..., we are a high‑conviction research team with the strategic patience... .... Define and evolve Evaluation Frameworks and Benchmarks for... ...transformers, MoE, large‑scale training). Generative world modeling...TrainingFull timeWork at officeWork from homeVisa sponsorshipRelocation packageFlexible hours
$69.76k
...Responsibilities: The Research Scientist I will work as part... ...outcomes to monitor and evaluate the effects of risk... ...specified in the job posting and must also be located... ...educational or training opportunities. Retirement... ...are welcome to apply. Work Location Expectations...TrainingContract workPart timeWork at officeLocal areaRemote workRelocationVisa sponsorshipFlexible hours$117.3k - $175.9k
# Battery Research ScientistOnsite |Mid Level|Full TimePosted... ...) within PSL conducts applied research in energy... ...for characterization, evaluation, and assessment in... ...As a Battery Research Scientist (Senior Member of Technical... ...which includes: training, motivating and directing...TrainingFull timeFor contractorsImmediate startRemote workWork visaRelocation packageFlexible hours$200k - $250k
...'ll have the compute power to train large models to solve this challenge. Your role as Applied Scientist is to build next-generation AI... ...fall short. Your focus: Applied research from research to data to model... ...and evals. Training (pre & post-training) and fine-tuning Large...TrainingRemote workFlexible hours$150k - $173k
...join a small but world-class Applied Research and AI team and work on... ...work in practice. Simulation & Evaluation: Build simulation environments... ...- and can be used to train, evaluate, and iterate on decision... ...monitoring and retraining RL systems post-deployment, including...TrainingBi-weekly payShift work$126k - $167k
...Research Scientist, Battlespace Awareness Anduril Industries is a defense... ...possess an M.S. or Ph.D. in Applied or Computational Mathematics... ...interview process in which we also evaluate practical experience and... ..., education and/or training, critical skills, and/or business...TrainingFull timeFor contractorsWork experience placementFor subcontractorFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!





