Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI/ML Research Engineer, LLM Training & Evaluation

Full-time

Innodata

Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model builders and leading labs.

This role is ideal for someone who has hands-on experience fine-tuning and evaluating large language models (and ideally multimodal models), and who can bridge research and engineering in real-world customer environments. You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client technical stakeholders to design and implement robust training/evaluation pipelines using both human-in-the-loop and AI-augmented methods.

The ideal candidate brings a strong computer science / machine learning engineering background, experience with modern LLM post-training workflows, and the ability to engage credibly with technical counterparts at leading AI organizations.

What You’ll Own:

As an AI/ML Research Engineer, LLM Training & Evaluation , you will design and implement the pipelines and tooling that connect data, evaluation, and post-training. You will help customers and internal teams move from evaluation findings to measurable model improvements.

Your work may include building fine-tuning workflows (e.g., supervised fine-tuning and preference-based optimization), integrating evaluation harnesses into model development loops, improving experiment reliability and throughput, and supporting advanced evaluation scenarios such as long-context, cross-modal, and dynamic multi-turn interactions.

You will also contribute to Innodata’s internal R&D efforts, including benchmark datasets, evaluation frameworks, and reusable infrastructure for model assessment and post-training experimentation. Additional responsibilities include (but are not limited to):

  • Lead or co-lead technically complex ML engineering projects from initial customer discussions through implementation and delivery
  • Design, build, and improve LLM training and post-training pipelines, including data ingestion, preprocessing, fine-tuning, evaluation, and experiment tracking
  • Implement and optimize evaluation systems for LLMs and multimodal models, including offline benchmarks and task-specific test harnesses
  • Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows
  • Build robust infrastructure and tooling for reproducible experimentation, metrics logging, and regression monitoring
  • Diagnose model behavior and pipeline failures, including data issues, training instability, metric inconsistencies, and evaluation drift
  • Collaborate with Language Data Scientists and Applied Research Scientists to translate evaluation frameworks into executable systems
  • Work closely with customer technical stakeholders to understand goals, constraints, and success criteria; propose and implement technically sound solutions
  • Contribute to internal research and platform development, including benchmark frameworks, evaluation tooling, and post-training workflow improvements
  • Contribute to best practices and standards for LLM training, evaluation, and quality assurance across projects
  • Mentor junior engineers and contribute to technical design reviews, documentation, and engineering rigor across the team

You’ll Thrive in This Role If You Have:

  • BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field ( MS/PhD preferred )
  • 2-3 years of relevant industry or research engineering experience in ML/AI systems
  • Hands-on experience with LLM training / fine-tuning / post-training, including at least one of:
    • supervised fine-tuning (SFT)
    • preference optimization (e.g., DPO or related methods)
    • RLHF / RLAIF-style workflows
    • task- or domain-adaptation of foundation models
  • Strong programming skills in Python and experience building production-quality ML code
  • Experience with modern ML frameworks (e.g., PyTorch, JAX, TensorFlow ) and model libraries/tooling (e.g., Hugging Face ecosystem, vLLM, distributed training stacks )
  • Experience designing and implementing evaluation pipelines for LLM/ML systems, including metrics computation, dataset handling, and experiment comparisons
  • Strong understanding of data pipelines and ML systems engineering, including reproducibility, observability, and debugging
  • Experience with large-scale distributed ML systems and performance optimization for training/evaluation workloads ( GPU/accelerator environments preferred )
  • Experience with large-scale data processing and workflow orchestration in support of model training/evaluation
  • Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data engineers, and customer technical leads
  • Strong written and verbal communication skills, including the ability to explain complex technical tradeoffs to both technical and non-technical audiences

Technical Skills

ML / LLM Engineering

  • Experience training, fine-tuning, and evaluating transformer-based models
  • Understanding of post-training workflows and model iteration loops
  • Familiarity with inference-time considerations (latency, throughput, memory/performance tradeoffs) where relevant to evaluation or deployment

Evaluation & Experimentation

  • Experience implementing automated evaluation pipelines and test harnesses
  • Experience with experiment tracking, versioning, and reproducibility practices
  • Ability to assess metric quality and ensure consistency across model comparisons

Software / Data Engineering

  • Proficiency in Python and strong software engineering fundamentals
  • Experience with data processing pipelines, storage formats, and scalable dataset workflows
  • Familiarity with CI/CD, testing, and engineering quality practices for ML systems
Vacancy posted 21 days ago
Similar jobs that could be interesting for youBased on the AI/ML Research Engineer, LLM Training & Evaluation in Remote vacancy
  •  ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model... 
    Training
    Full time

    Innodata

    Remote
    9 days ago
  •  ...full-time opportunity for machine learning engineers and research practitioners to design, implement, and evaluate end-to-end ML benchmarks. You will transform real ML ideas...  ...experiments in Python notebooks, analyze training behavior, and assess model-generated... 
    Training
    Hourly pay
    Full time
    Remote work

    24-Mag Llc

    New York, NY
    2 days ago
  • $213k - $263k

     ...The mission of the Waymo AI Foundations team is to...  ...with other research teams in Alphabet. AI Foundations...  ...hierarchical learning, and robust evaluation. This role follows a...  ...models. Explore LLM/VLM distillation...  ...prefer: Experience in training or deploying multi-... 
    Training
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  • $238k - $302k

     ...mission of the Waymo AI Foundations team...  ...with other research teams in Alphabet....  ...learning, and robust evaluation. This role follows...  ...Senior Staff Software Engineer.   You will:...  ...of experience in ML engineering and...  ...infra experience: training, evaluating and deploying... 
    Training
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  • $281k - $356k

     ...and speed up the evaluation and onboard developer...  ...to deliver training and evaluation data...  ...are looking for researchers and software engineers who are passionate...  ...across our embodied AI applications....  ...and Generative AI (LLM/VLM) solutions....  ...Python and standard ML frameworks (e.g.,... 
    Training
    Full time

    Waymo

    Remote
    1 day ago
  • $230k - $322k

    Role Description The AI Engineering team at Reddit is...  ...of applied research and massive-scale infrastructure, training models that truly...  ...for Post-Training & Evaluation Science, you will...  ...years of professional ML experience (or PhD...  ...a direct focus on LLM post-training and... 
    Training
    Full time
    Flexible hours

    Reddit

    Remote
    5 days ago
  • $216k - $270k

     ...develop reliable AI systems for the world...  ...rigorous evaluation with full-stack deployment...  ...with applied ML research, design, and evaluation...  ...Learning Research Engineer, you will operate...  .... This could mean training and fine-tuning models...  ...training methods, LLM alignment, or... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  • $128k - $159.95k

    Calabasas, CAScience & Engineering - Intelligent...  ...Lead and conduct research in agentic AI, intelligent...  ...autonomous workflows, and LLM-powered agent...  ..., develop, and evaluate multi-agent systems...  ...experience in AI/ML• Strong...  ...relevant education or training. Your recruitercan... 
    Training
    Full time

    HRL Laboratories

    Calabasas, CA
    4 days ago
  • $250.3k - $289k

     ...Description As a Principal AI/ML Research Engineer, this technical leader will...  ...Design self-supervised pre-training strategies on raw payment histories...  ...AI frameworks, including LLM fine-tuning, retrieval-...  ...of effective AI agent evaluation methodologies. ~Design and... 
    Training
    Flexible hours

    WEXWEXUS

    Remote
    1 day ago
  • $75 - $90 per hour

     ...technical talent with leading AI research labs. Headquartered in San...  ...Human Baseliner for Open-Ended ML Research Tasks Type: Contract...  ...completed questionnaires for evaluation. Qualifications Must-...  ...design, contrastive training, generative modeling, multilingual... 
    Training
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    a month ago
  • $150k - $230k

     ...powered by advanced AI, recommendation...  ...Learning Engineer to drive the post-training of our large language...  ...and maintain evaluation and reward/verifier...  ...post-training research and turn...  ...Hands-on LLM post-training experience...  ...engineering for ML. You can independently... 
    Training
    Full time
    Local area
    Work from home

    News Break

    Remote
    1 day ago
  •  ...is the Agentic AI Assistant platform...  ...’ Reasoning Engine and natural language...  ...cutting edge ML infrastructure for...  ...building and serving LLM’s at Moveworks....  ...distributed training and inference pipeline...  ...(LLM), model evaluation and monitoring...  .... ~ A love of research publications in... 
    Training
    Full time
    Work at office
    Remote work
    Flexible hours

    Servicenow

    Remote
    1 day ago
  •  ...entire robotics stack. We're training state-of-the-art AI models that leverage our...  ...As a Machine Learning Research Engineer, you will work on the software...  ...to model training, evaluation, and on-robot deployment...  ...model training pipelines (LLM/VLM/VLA) At Sunday... 
    Training

    Sunday

    Redwood City, CA
    3 days ago
  •  ...operate production-grade LLM-based systems that...  ...expertise into robust AI products. You will design...  .... This role is core engineering work on a remote, high...  ...impact team focused on training and evaluating frontier AI models....  ...build, and deploy AI and ML solutions with... 
    Training
    Full time
    Remote work

    SaidGig

    United States
    14 days ago
  •  ...Marketing we rely on ML to ensure that...  ...Airbnb.  The CS AI product team is responsible...  ...tools including LLM fine-tuning,...  ...optimization, RAG/Search, LLM evaluation and testing...  ...principal machine learning engineer, you will be...  ...~ Proven record of training, fine tuning, optimizing... 
    Training
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    United States
    1 day ago
  • $189.4k - $300.6k

     ...driving? Join the Embodied AI team at General Motors....  ...-world scenarios.The Evaluation Foundations team—part...  ...vehicle development. We engineer high-performance tools...  ...partner with data-intensive ML teams to drive rapid...  ...to data selection, training, and launch decisions.Drive... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    3 days ago
  • $2,000 per month

     ...Elastic, the Search AI Company, enables everyone...  ...a Principal Security ML Research Engineer on the Threat Research...  ...of ML models. Build evaluation frameworks to assess...  ...algorithms, and experience training models using scikit-...  ...integrating LLM APIs into production applications... 
    Training
    Full time
    Local area
    Flexible hours

    Elastic

    Remote
    14 days ago
  • $85 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by completing...  ...machine learning engineering tasks and...  ...reflect production ML workflows, helping a leading AI research lab identify...  ...involving model training, inference systems, MLOps, and LLM applications.... 
    Training
    Hourly pay
    Remote work

    SaidGig

    United States
    16 days ago
  • $165k - $310k

     ...Are Lightning AI is the company behind...  ...for developing, training, and deploying AI...  ...to take ideas from research to production with...  ...Senior Research Engineer who has built, trained...  ..., fine-tuned, evaluated, and deployed across...  ...with modern LLM training and post-... 
    Training
    Full time
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    Seattle, WA
    2 days ago
  • $189.4k - $300.6k

     ...driving? Join the Embodied AI team at General Motors....  ...world scenarios. The Evaluation Foundations team—part...  ...development. We engineer high-performance tools...  ...partner with data-intensive ML teams to drive rapid innovation...  ...to data selection, training, and launch decisions.... 
    Training
    Full time
    Relocation package
    Flexible hours

    General Motors

    Remote
    1 day ago
  •  ...As a full-time remote ML Research Engineer focused on AI for Life Sciences, the successful candidate will bring cutting-edge research into production...  ...software in professional teams Experience or training in data-science tasks related to structural biology Expertise... 
    Training
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  •  ...Job Responsibilities: Engineer, design, implement, and improve...  ...systems and tools for enabling research Apply knowledge of relevant...  ...~1-2 years of Distributed ML Training (FSDP/DDP) experience ~1-3...  ...contributions to open-source AI/ML projects Education/Experience... 
    Training
    Work experience placement

    SGS Consulting

    Remote
    more than 2 months ago
  •  ...Job Responsibilities: Engineer, design, implement, and improve...  ...systems and tools for enabling research Apply knowledge of relevant...  ...~5+ years of Distributed ML Training (FSDP/DDP) experience ~3-5...  ...contributions to open-source AI/ML projects ~ ears of Dataset... 
    Training
    Work experience placement

    SGS Consulting

    Remote
    more than 2 months ago
  •  ...growth company delivering AI solutions that address...  ...mathematics, medicine, engineering, and other specialties....  ...creates leading edge ML and physics-based models...  ...medicines. As a ML Research Engineer, you will bring...  ...distributed training pipelines on world-class... 
    Training
    Full time
    Seasonal work
    Flexible hours

    SandboxAQ

    Remote
    28 days ago
  • $100k - $150k

     ...Description We are seeking an AI Research Engineer to bridge cutting-edge...  ...deep understanding of modern ML and deep learning techniques...  ...research landscape, can critically evaluate new techniques for real-...  ...~Build production-quality training and inference pipelines using... 
    Training
    Full time
    Local area
    Immediate start

    Bright Vision Technologies

    Remote
    1 day ago
  • $213k - $263k

     ...dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge:...  ...We are seeking visionary machine learning engineers and researchers to architect the scalable deep learning systems... 
    Training
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $204k - $259k

     ...Driver. We conduct our own research to address real-world...  ...of sensors, enabling engineers like you to (1)...  ...develop models and model training at scale, to (3) analyze...  ...and rigorously evaluate metrics and methodologies...  ...scale model development (LLM, VLM, or similar foundation... 
    Training
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $174k - $253k

     ...and implement robust agents and LLM-powered journeys that evaluate, gate, and improve engineering artifacts.Engineer and refine skills...  ...and context provided to the AI agents, enhancing their capabilities...  ..., and relevant education or training. US: $174000 - $253000 (USD) + 1... 
    Training

    Google

    San Jose, CA
    3 days ago
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered in...  .... Position: ML Engineer (Coding Agent...  ...agents to complete and evaluate complex machine learning...  ...implementations involving model training , inference systems , MLOps , and LLM applications .... 
    Training
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 day ago
  • $204k - $259k

     ...Driver Understanding and Evaluation (DUE) team at Waymo...  ...models to deliver training and evaluation data...  ...We are looking for researchers and software engineers who are passionate about...  ...scalable and robust ML pipelines for...  ...modeling and generative AI into robust,... 
    Training
    Full time

    Waymo

    Remote
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI/ML Research Engineer, LLM Training & Evaluation. Be the first to apply!