Applied Research Scientist, LLM Evaluation & Post-Training
$175k - $225kInnodata Inc.
Role Description
Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.
This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.
The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.
What You’ll Own
- Define the next generation of evaluation-driven model improvement workflows.
- Study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes.
- Design experiments that produce credible, actionable conclusions.
- Design benchmark datasets, develop evaluation taxonomies and protocols, define metrics and scoring methodologies, analyze failure modes, and test how changes in evaluation setup affect downstream fine-tuning results.
- Support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
- Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement.
- Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes.
- Develop and validate evaluation frameworks for LLM and multimodal systems, including:
- benchmark/task design
- scoring methods
- judge/model-assisted evaluation
- human evaluation protocols
- robustness/stress testing
- Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations.
- Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs.
- Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign.
- Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines.
- Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs.
- Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations.
- Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets.
- Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations.
- Contribute to thought leadership and best practices in LLM evaluation, post-training, and GenAI quality measurement.
Qualifications
- MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field (PhD strongly preferred).
- 5+ years of relevant experience in applied research / research science in ML/AI, with substantial work in LLMs or foundation models.
- Demonstrated experience with LLM evaluation, benchmarking, alignment, post-training, or model quality research.
- Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems.
- Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization).
- Experience working with modern ML tooling/frameworks (e.g., PyTorch, Hugging Face, JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments.
- Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability.
- Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs.
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts.
- Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly.
Requirements
- The expected salary range for this position is $175,000 – $225,000 USD per year, based on experience, skills, and qualifications.
Company Description
Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.
- Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training , you will lead research and experimentation on how evaluation...TrainingFull time
$142.8k - $193.2k
...individual customers, building evaluation frameworks for model... ...designing data‑driven guardrails for LLM‑generated content. The work... ...architecture design and deep learning training and optimization and model... ...If the country/region you’re applying in isn’t listed, please...TrainingLocal areaWorldwideFlexible hours$171.6k - $222.2k
...You will design, train, and optimize generative... ...involves building LLM-based classifiers... ...data engines, and post-generation content... ...metrics, defining evaluation frameworks that... ...-knit group of scientists and engineers who... ...country/region you’re applying in isn’t listed,...TrainingLocal areaFlexible hours$167.1k - $226.1k
...scalable and reliable evaluation of state-of-the-art Conversational... ..., and resourceful Applied Scientist in the field of Large... ...with LLMs, including LLM‑as‑a‑Judge (LLMaaJ),... ...expertise to set the research agenda for how we... ...sources, and design, train, and maintain the evaluation...TrainingLocal areaFlexible hours$126.2k - $264.1k
...As our Principal Applied Scientist, you will play a key... ...drive projects from research POC to production.... ...including data, model, training, and evaluation, employing best practices... ...technologies in LLM and generative AI, such... ...provided in this posting are specific to the...TrainingTemporary workRemote workFlexible hours$100 - $120 per hour
...LLM Research Scientist (Pre-training & Post-Training) is a remote review track for evaluating AI outputs across research science reasoning, calculations, and research workflows.... ...Graduate-level training or equivalent applied experience in research science or a closely...TrainingRemote jobFor contractors10 hours per week$167.1k - $226.1k
...responsibilities As an Applied Scientist on LPEX, you will be a technical... ..., translating research advances into measurable business... ...development, offline and online evaluation, and reliable production... ...architecture design and deep learning training and optimization and model...TrainingWorldwideFlexible hoursNight shift$167.1k - $226.1k
...innovative and customer-focused applied scientist to help us make the world's... ...at the frontier of AI research, and rapidly implement and... ...multimodal data with rigorous evaluation frameworks. Define research... ...Experience with training and deploying machine learning...TrainingWorldwideFlexible hours$149k - $350k
...join us! We’re looking for applied scientists with a Machine Learning and... ...fundamental and applied research in this area. You will be combining... ...for AI systems Build evaluation systems to measure and... ...HuggingFace etc ~ Experience training LLMs with Reinforcement...TrainingFull timeTemporary workRemote workWork from home- ...business problems. We’re training and deploying... ...is a team of researchers, engineers, designers... ...us! Why this role? Evaluation is critical to... ...infrastructure to measure LLM progress. As a Senior Research Scientist, Model Evaluation,... ...encourage you to apply. We strive to...TrainingFull timeWork at officeLocal areaRemote workHome office
- ...The Role We're looking for an Applied Scientist who thrives at the intersection of applied research and real-world products. You'll... ...in data, interaction, and evaluation with both creativity and engineering... ...of data modeling for training and how curation decisions shape...TrainingFlexible hours
- ...power grids, buildings, trains, hospitals. Industrial... ...major opportunity in applied AI, and one of the... ...intersection of machine learning research, real world data, and... .... As Senior Applied Scientist, you lead the science... ...control, planning, or evaluation Take problems from...Training
- ...focused, healthcare‑only LLM platform enables... ...voice. As a Senior Staff Research Scientist in Speech Technologies... ...datasets, creating the training foundation that gives... ...medical domain Train, evaluate, and optimize ASR models... ...to match. Ready to Apply? If you've spent your...TrainingWork at officeRemote work
$80 - $110 per hour
...expertise to the development and evaluation of next-generation AI systems that must reason about pure and applied mathematics at the level of a working research mathematician. You will help ensure... ...research-level mathematics problems to train and evaluate frontier AI models....TrainingHourly payPart timeImmediate startRemote work$80 - $110 per hour
...Contribute frontier research expertise to the development and evaluation of AI systems that... ...supports model training and assessment across... ...member, principal scientist, or industry... ...are encouraged to apply, the depth of research... ...application through the posting you are viewing,...TrainingHourly payPart timeImmediate startRemote work- ...skilled and driven Senior Applied Scientist with 7+ years of industry experience... .... Responsibilities ~Research and engineer NLP solutions... ...and implement algorithms, train state of the art large language models (LLM) on large data, and evaluate their performance. ~...TrainingFull time
$88.6k - $150.8k
Join to apply for the Research Scientist role at The Henry M. Jackson Foundation for... ...clinical tool development and training, cognitive monitoring and... ...on employee performance evaluations. Qualifications Education... ...Learning Scientist, NLP/LLM San Francisco, CA $140,000...TrainingFull timeContract workFor contractorsFixed term contractInternshipWork at officeLocal areaRemote workFlexible hours$40 per hour
...We are looking for an Applied Mathematician to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level of...Hourly payFull timeContract workPart timeRemote work$150k - $173k
...join a small but world-class Applied Research and AI team and work on... ...in practice. * Simulation & Evaluation: Build simulation environments... ...operations — and can be used to train, evaluate, and iterate on... ...and retraining RL systems post-deployment, including...TrainingBi-weekly payFull timeShift work$150k - $230k
...Engineer to drive the post-training of our large language... ...large GPU clusters , applying distributed-training... ...Build and maintain evaluation and reward/verifier pipelines... ...with post-training research and turn promising... ...Requirements Hands-on LLM post-training...TrainingFull timeLocal areaWork from home$200k - $270k
...looking for a Staff Research Scientist to join Cognitiv’... ...the AdTech and LLM landscape well... ...embeddings, and applied deep learning to... ...including large‑scale training, fine‑tuning, and... ...of a launch—blog posts, talks, demos, or... ...candidate evaluation or hiring decisions...TrainingWork at officeRemote workWork from home$75 per hour
...Role Overview AI & Machine Learning Researchers evaluate AI-generated content and create expert-level training data to improve AI systems'' understanding of advanced computer... ...(cs), and the ability to interpret and apply findings from recent preprints. Strong written...TrainingHourly payContract workPart timeFor contractorsRemote workFlexible hours- A leading data services company is seeking an Applied Mathematician to evaluate AI models by providing complex mathematical problems to chatbots and assessing their outputs for quality and performance. This role offers flexibility with fully remote work and allows you to...Hourly payRemote work
$311.85k - $370k
...Wayve AI Platform Scientist Or Engineer Founded in... ...we curate, enrich, and evaluate the real-world driving... ...are hiring at either Applied Scientist or Machine Learning... ...into high-signal training data through (semi-)... ...record of taking ML from research into production...TrainingFull timeWork at officeWork from home- ...growth. If you are ready to apply your skills to the... ...robotics stack. We're training state-of-the-art AI models... ...As a Machine Learning Research Engineer, you will work... ...to model training, evaluation, and on-robot deployment... ...model training pipelines (LLM/VLM/VLA) At...Training
$200k - $250k
...'ll have the compute power to train large models to solve this challenge. Your role as Applied Scientist is to build next-generation AI... ...fall short. Your focus: Applied research from research to data to model... ...and evals. Training (pre & post-training) and fine-tuning Large...TrainingRemote workFlexible hours$183.8k - $248.7k
...the full ML lifecycle, from exploratory research and offline modeling to online experimentation... ...analyze large-scale A/B experiments, applying causal inference techniques to measure... ...fundamentals, including architecture, training/inference lifecycles, and optimization of...TrainingFlexible hoursShift work$30 per hour
A leading AI training firm is seeking an Applied Physics Research Scientist to evaluate AI chatbots on physics-related challenges. This is a remote position ideal for experts with a strong grasp of classical mechanics and related fields. Applicants must possess fluency...TrainingHourly payRemote workFlexible hours$125k - $225k
...FutureSearch is looking for exceptional Research Scientists to evaluate and improve state-of-the-art forecasting and agentic LLM web research. We are an elite team of engineers... ...frontier labs on research, evaluation, and training. You are a talented researcher with...TrainingRemote workFlexible hours$290.25k
...Machine Learning Scientists to join our AI team... ...AI applications (LLM and Computer Vision... ...key member of our research and development efforts... ...architectures. Evaluate the performance of... ...architectures, train/evaluate/tune models... ...less likely to apply to jobs unless they...TrainingWork experience placementWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied Research Scientist, LLM Evaluation & Post-Training. Be the first to apply!






