AI/ML Research Engineer, LLM Training & Evaluation
Innodata
Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model builders and leading labs.
This role is ideal for someone who has hands-on experience fine-tuning and evaluating large language models (and ideally multimodal models), and who can bridge research and engineering in real-world customer environments. You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client technical stakeholders to design and implement robust training/evaluation pipelines using both human-in-the-loop and AI-augmented methods.
The ideal candidate brings a strong computer science / machine learning engineering background, experience with modern LLM post-training workflows, and the ability to engage credibly with technical counterparts at leading AI organizations.
What You’ll Own:
As an AI/ML Research Engineer, LLM Training & Evaluation , you will design and implement the pipelines and tooling that connect data, evaluation, and post-training. You will help customers and internal teams move from evaluation findings to measurable model improvements.
Your work may include building fine-tuning workflows (e.g., supervised fine-tuning and preference-based optimization), integrating evaluation harnesses into model development loops, improving experiment reliability and throughput, and supporting advanced evaluation scenarios such as long-context, cross-modal, and dynamic multi-turn interactions.
You will also contribute to Innodata’s internal R&D efforts, including benchmark datasets, evaluation frameworks, and reusable infrastructure for model assessment and post-training experimentation. Additional responsibilities include (but are not limited to):
- Lead or co-lead technically complex ML engineering projects from initial customer discussions through implementation and delivery
- Design, build, and improve LLM training and post-training pipelines, including data ingestion, preprocessing, fine-tuning, evaluation, and experiment tracking
- Implement and optimize evaluation systems for LLMs and multimodal models, including offline benchmarks and task-specific test harnesses
- Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows
- Build robust infrastructure and tooling for reproducible experimentation, metrics logging, and regression monitoring
- Diagnose model behavior and pipeline failures, including data issues, training instability, metric inconsistencies, and evaluation drift
- Collaborate with Language Data Scientists and Applied Research Scientists to translate evaluation frameworks into executable systems
- Work closely with customer technical stakeholders to understand goals, constraints, and success criteria; propose and implement technically sound solutions
- Contribute to internal research and platform development, including benchmark frameworks, evaluation tooling, and post-training workflow improvements
- Contribute to best practices and standards for LLM training, evaluation, and quality assurance across projects
- Mentor junior engineers and contribute to technical design reviews, documentation, and engineering rigor across the team
You’ll Thrive in This Role If You Have:
- BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field ( MS/PhD preferred )
- 2-3 years of relevant industry or research engineering experience in ML/AI systems
- Hands-on experience with LLM training / fine-tuning / post-training, including at least one of:
- supervised fine-tuning (SFT)
- preference optimization (e.g., DPO or related methods)
- RLHF / RLAIF-style workflows
- task- or domain-adaptation of foundation models
- Strong programming skills in Python and experience building production-quality ML code
- Experience with modern ML frameworks (e.g., PyTorch, JAX, TensorFlow ) and model libraries/tooling (e.g., Hugging Face ecosystem, vLLM, distributed training stacks )
- Experience designing and implementing evaluation pipelines for LLM/ML systems, including metrics computation, dataset handling, and experiment comparisons
- Strong understanding of data pipelines and ML systems engineering, including reproducibility, observability, and debugging
- Experience with large-scale distributed ML systems and performance optimization for training/evaluation workloads ( GPU/accelerator environments preferred )
- Experience with large-scale data processing and workflow orchestration in support of model training/evaluation
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data engineers, and customer technical leads
- Strong written and verbal communication skills, including the ability to explain complex technical tradeoffs to both technical and non-technical audiences
Technical Skills
ML / LLM Engineering
- Experience training, fine-tuning, and evaluating transformer-based models
- Understanding of post-training workflows and model iteration loops
- Familiarity with inference-time considerations (latency, throughput, memory/performance tradeoffs) where relevant to evaluation or deployment
Evaluation & Experimentation
- Experience implementing automated evaluation pipelines and test harnesses
- Experience with experiment tracking, versioning, and reproducibility practices
- Ability to assess metric quality and ensure consistency across model comparisons
Software / Data Engineering
- Proficiency in Python and strong software engineering fundamentals
- Experience with data processing pipelines, storage formats, and scalable dataset workflows
- Familiarity with CI/CD, testing, and engineering quality practices for ML systems
- ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model...TrainingFull time
- ...full-time opportunity for machine learning engineers and research practitioners to design, implement, and evaluate end-to-end ML benchmarks. You will transform real ML ideas... ...experiments in Python notebooks, analyze training behavior, and assess model-generated...TrainingHourly payFull timeRemote work
$213k - $263k
...The mission of the Waymo AI Foundations team is to... ...with other research teams in Alphabet. AI Foundations... ...hierarchical learning, and robust evaluation. This role follows a... ...models. Explore LLM/VLM distillation... ...prefer: Experience in training or deploying multi-...TrainingFull timeRemote work$238k - $302k
...mission of the Waymo AI Foundations team... ...with other research teams in Alphabet.... ...learning, and robust evaluation. This role follows... ...Senior Staff Software Engineer. You will:... ...of experience in ML engineering and... ...infra experience: training, evaluating and deploying...TrainingFull timeRemote work$281k - $356k
...and speed up the evaluation and onboard developer... ...to deliver training and evaluation data... ...are looking for researchers and software engineers who are passionate... ...across our embodied AI applications.... ...and Generative AI (LLM/VLM) solutions.... ...Python and standard ML frameworks (e.g.,...TrainingFull time$230k - $322k
Role Description The AI Engineering team at Reddit is... ...of applied research and massive-scale infrastructure, training models that truly... ...for Post-Training & Evaluation Science, you will... ...years of professional ML experience (or PhD... ...a direct focus on LLM post-training and...TrainingFull timeFlexible hours$216k - $270k
...develop reliable AI systems for the world... ...rigorous evaluation with full-stack deployment... ...with applied ML research, design, and evaluation... ...Learning Research Engineer, you will operate... .... This could mean training and fine-tuning models... ...training methods, LLM alignment, or...TrainingFull time$128k - $159.95k
Calabasas, CAScience & Engineering - Intelligent... ...Lead and conduct research in agentic AI, intelligent... ...autonomous workflows, and LLM-powered agent... ..., develop, and evaluate multi-agent systems... ...experience in AI/ML• Strong... ...relevant education or training. Your recruitercan...TrainingFull time$250.3k - $289k
...Description As a Principal AI/ML Research Engineer, this technical leader will... ...Design self-supervised pre-training strategies on raw payment histories... ...AI frameworks, including LLM fine-tuning, retrieval-... ...of effective AI agent evaluation methodologies. ~Design and...TrainingFlexible hours$75 - $90 per hour
...technical talent with leading AI research labs. Headquartered in San... ...Human Baseliner for Open-Ended ML Research Tasks Type: Contract... ...completed questionnaires for evaluation. Qualifications Must-... ...design, contrastive training, generative modeling, multilingual...TrainingContract workFor contractorsSummer workRemote work$150k - $230k
...powered by advanced AI, recommendation... ...Learning Engineer to drive the post-training of our large language... ...and maintain evaluation and reward/verifier... ...post-training research and turn... ...Hands-on LLM post-training experience... ...engineering for ML. You can independently...TrainingFull timeLocal areaWork from home- ...is the Agentic AI Assistant platform... ...’ Reasoning Engine and natural language... ...cutting edge ML infrastructure for... ...building and serving LLM’s at Moveworks.... ...distributed training and inference pipeline... ...(LLM), model evaluation and monitoring... .... ~ A love of research publications in...TrainingFull timeWork at officeRemote workFlexible hours
- ...entire robotics stack. We're training state-of-the-art AI models that leverage our... ...As a Machine Learning Research Engineer, you will work on the software... ...to model training, evaluation, and on-robot deployment... ...model training pipelines (LLM/VLM/VLA) At Sunday...Training
- ...operate production-grade LLM-based systems that... ...expertise into robust AI products. You will design... .... This role is core engineering work on a remote, high... ...impact team focused on training and evaluating frontier AI models.... ...build, and deploy AI and ML solutions with...TrainingFull timeRemote work
- ...Marketing we rely on ML to ensure that... ...Airbnb. The CS AI product team is responsible... ...tools including LLM fine-tuning,... ...optimization, RAG/Search, LLM evaluation and testing... ...principal machine learning engineer, you will be... ...~ Proven record of training, fine tuning, optimizing...TrainingRemote jobFull timeCasual workLive inWork at office
$189.4k - $300.6k
...driving? Join the Embodied AI team at General Motors.... ...-world scenarios.The Evaluation Foundations team—part... ...vehicle development. We engineer high-performance tools... ...partner with data-intensive ML teams to drive rapid... ...to data selection, training, and launch decisions.Drive...TrainingFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$2,000 per month
...Elastic, the Search AI Company, enables everyone... ...a Principal Security ML Research Engineer on the Threat Research... ...of ML models. Build evaluation frameworks to assess... ...algorithms, and experience training models using scikit-... ...integrating LLM APIs into production applications...TrainingFull timeLocal areaFlexible hours$85 per hour
...Role Overview Evaluate and improve frontier AI coding agents by completing... ...machine learning engineering tasks and... ...reflect production ML workflows, helping a leading AI research lab identify... ...involving model training, inference systems, MLOps, and LLM applications....TrainingHourly payRemote work$165k - $310k
...Are Lightning AI is the company behind... ...for developing, training, and deploying AI... ...to take ideas from research to production with... ...Senior Research Engineer who has built, trained... ..., fine-tuned, evaluated, and deployed across... ...with modern LLM training and post-...TrainingFull timeWork at officeRemote workWork from homeFlexible hours2 days per week$189.4k - $300.6k
...driving? Join the Embodied AI team at General Motors.... ...world scenarios. The Evaluation Foundations team—part... ...development. We engineer high-performance tools... ...partner with data-intensive ML teams to drive rapid innovation... ...to data selection, training, and launch decisions....TrainingFull timeRelocation packageFlexible hours- ...As a full-time remote ML Research Engineer focused on AI for Life Sciences, the successful candidate will bring cutting-edge research into production... ...software in professional teams Experience or training in data-science tasks related to structural biology Expertise...TrainingFull timeRemote work
- ...Job Responsibilities: Engineer, design, implement, and improve... ...systems and tools for enabling research Apply knowledge of relevant... ...~1-2 years of Distributed ML Training (FSDP/DDP) experience ~1-3... ...contributions to open-source AI/ML projects Education/Experience...TrainingWork experience placement
- ...Job Responsibilities: Engineer, design, implement, and improve... ...systems and tools for enabling research Apply knowledge of relevant... ...~5+ years of Distributed ML Training (FSDP/DDP) experience ~3-5... ...contributions to open-source AI/ML projects ~ ears of Dataset...TrainingWork experience placement
- ...growth company delivering AI solutions that address... ...mathematics, medicine, engineering, and other specialties.... ...creates leading edge ML and physics-based models... ...medicines. As a ML Research Engineer, you will bring... ...distributed training pipelines on world-class...TrainingFull timeSeasonal workFlexible hours
$100k - $150k
...Description We are seeking an AI Research Engineer to bridge cutting-edge... ...deep understanding of modern ML and deep learning techniques... ...research landscape, can critically evaluate new techniques for real-... ...~Build production-quality training and inference pipelines using...TrainingFull timeLocal areaImmediate start$213k - $263k
...dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge:... ...We are seeking visionary machine learning engineers and researchers to architect the scalable deep learning systems...TrainingFull timeRemote work$204k - $259k
...Driver. We conduct our own research to address real-world... ...of sensors, enabling engineers like you to (1)... ...develop models and model training at scale, to (3) analyze... ...and rigorously evaluate metrics and methodologies... ...scale model development (LLM, VLM, or similar foundation...TrainingFull timeRemote work$174k - $253k
...and implement robust agents and LLM-powered journeys that evaluate, gate, and improve engineering artifacts.Engineer and refine skills... ...and context provided to the AI agents, enhancing their capabilities... ..., and relevant education or training. US: $174000 - $253000 (USD) + 1...Training$85 per hour
...technical talent with leading AI research labs. Headquartered in... .... Position: ML Engineer (Coding Agent... ...agents to complete and evaluate complex machine learning... ...implementations involving model training , inference systems , MLOps , and LLM applications ....TrainingFull timeContract workSummer workRemote work$204k - $259k
...Driver Understanding and Evaluation (DUE) team at Waymo... ...models to deliver training and evaluation data... ...We are looking for researchers and software engineers who are passionate about... ...scalable and robust ML pipelines for... ...modeling and generative AI into robust,...TrainingFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI/ML Research Engineer, LLM Training & Evaluation. Be the first to apply!






