AI/ML Research Engineer, LLM Training & Evaluation
Innodata
Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model builders and leading labs.
This role is ideal for someone who has hands-on experience fine-tuning and evaluating large language models (and ideally multimodal models), and who can bridge research and engineering in real-world customer environments. You will work closely with Language Data Scientists, Applied Research Scientists, data engineers, and client technical stakeholders to design and implement robust training/evaluation pipelines using both human-in-the-loop and AI-augmented methods.
The ideal candidate brings a strong computer science / machine learning engineering background, experience with modern LLM post-training workflows, and the ability to engage credibly with technical counterparts at leading AI organizations.
What You’ll Own:
As an AI/ML Research Engineer, LLM Training & Evaluation , you will design and implement the pipelines and tooling that connect data, evaluation, and post-training. You will help customers and internal teams move from evaluation findings to measurable model improvements.
Your work may include building fine-tuning workflows (e.g., supervised fine-tuning and preference-based optimization), integrating evaluation harnesses into model development loops, improving experiment reliability and throughput, and supporting advanced evaluation scenarios such as long-context, cross-modal, and dynamic multi-turn interactions.
You will also contribute to Innodata’s internal R&D efforts, including benchmark datasets, evaluation frameworks, and reusable infrastructure for model assessment and post-training experimentation. Additional responsibilities include (but are not limited to):
- Lead or co-lead technically complex ML engineering projects from initial customer discussions through implementation and delivery
- Design, build, and improve LLM training and post-training pipelines, including data ingestion, preprocessing, fine-tuning, evaluation, and experiment tracking
- Implement and optimize evaluation systems for LLMs and multimodal models, including offline benchmarks and task-specific test harnesses
- Integrate human-in-the-loop and AI-augmented evaluation signals into model development workflows
- Build robust infrastructure and tooling for reproducible experimentation, metrics logging, and regression monitoring
- Diagnose model behavior and pipeline failures, including data issues, training instability, metric inconsistencies, and evaluation drift
- Collaborate with Language Data Scientists and Applied Research Scientists to translate evaluation frameworks into executable systems
- Work closely with customer technical stakeholders to understand goals, constraints, and success criteria; propose and implement technically sound solutions
- Contribute to internal research and platform development, including benchmark frameworks, evaluation tooling, and post-training workflow improvements
- Contribute to best practices and standards for LLM training, evaluation, and quality assurance across projects
- Mentor junior engineers and contribute to technical design reviews, documentation, and engineering rigor across the team
You’ll Thrive in This Role If You Have:
- BS/MS/PhD in Computer Science, Machine Learning, AI, Applied Mathematics, or a related quantitative technical field ( MS/PhD preferred )
- 2-3 years of relevant industry or research engineering experience in ML/AI systems
- Hands-on experience with LLM training / fine-tuning / post-training, including at least one of:
- supervised fine-tuning (SFT)
- preference optimization (e.g., DPO or related methods)
- RLHF / RLAIF-style workflows
- task- or domain-adaptation of foundation models
- Strong programming skills in Python and experience building production-quality ML code
- Experience with modern ML frameworks (e.g., PyTorch, JAX, TensorFlow ) and model libraries/tooling (e.g., Hugging Face ecosystem, vLLM, distributed training stacks )
- Experience designing and implementing evaluation pipelines for LLM/ML systems, including metrics computation, dataset handling, and experiment comparisons
- Strong understanding of data pipelines and ML systems engineering, including reproducibility, observability, and debugging
- Experience with large-scale distributed ML systems and performance optimization for training/evaluation workloads ( GPU/accelerator environments preferred )
- Experience with large-scale data processing and workflow orchestration in support of model training/evaluation
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data engineers, and customer technical leads
- Strong written and verbal communication skills, including the ability to explain complex technical tradeoffs to both technical and non-technical audiences
Technical Skills
ML / LLM Engineering
- Experience training, fine-tuning, and evaluating transformer-based models
- Understanding of post-training workflows and model iteration loops
- Familiarity with inference-time considerations (latency, throughput, memory/performance tradeoffs) where relevant to evaluation or deployment
Evaluation & Experimentation
- Experience implementing automated evaluation pipelines and test harnesses
- Experience with experiment tracking, versioning, and reproducibility practices
- Ability to assess metric quality and ensure consistency across model comparisons
Software / Data Engineering
- Proficiency in Python and strong software engineering fundamentals
- Experience with data processing pipelines, storage formats, and scalable dataset workflows
- Familiarity with CI/CD, testing, and engineering quality practices for ML systems
$100 - $175 per hour
...Overview Join a cutting-edge AI evaluation program as a senior digital chip design and verification engineer, focusing on frontier silicon... ...analysis ~ Experience with LLM-based tools to enhance workflows... ...APB) and experience in CPU, GPU/ML accelerator, networking,...SuggestedRemote jobHourly payFull time- ...closely with experienced ML engineers, platform partners,... ...feature pipelines and training datasets from proprietary... ...define requirements, evaluate tradeoffs, and... ...have 6+ years experience researching, training, tuning and... ...Proficient in using AI-powered developer tools...TrainingFull time
- ...domains for generative AI. Numerical... ...Scientist, Financial AI Evaluation & Datasets , you... ...datasets used to train, fine-tune,... ...engagement), an Applied Research Scientist (shapes... ...), an AI/ML Research Engineer (builds training... ...— rubric-grounded LLM-as-judge prompts,...TrainingFull timeShift work
- ...a global data engineering company. We believe... ...Intelligence (AI) are... ...providing the data, evaluation frameworks, and... ...datasets used to train, fine-tune, and... ...data science, and ML stakeholders.... ...Architect, Applied Research Scientist, AI/... ...-grounded LLM-as-judge prompts...TrainingFull timeShift work
$200k - $300k
...Machine Learning Engineer, you will design,... ...learning and generative AI systems that power... ...advanced ML and LLM capabilities into... ...than purely academic research. Responsibilities... ...business stakeholders to evaluate buy vs. build... ..., model training, evaluation, deployment...TrainingFull timeWork at office$164.49k - $197.39k
...Cloud's actually useful AI, organizations can see,... ...-based recommendation engine that provides useful contextual... .... We are hiring an ML Engineer to lead its... ...experiments, establish evaluation methodology, and define... .... You’ll own model training & serving Define what...TrainingFull timeLocal areaRemote workFlexible hours$50 per hour
...Role Overview As a Software Engineering evaluator, you will play a crucial... ...creating advanced datasets for training, benchmarking, and enhancing... ...collaborating closely with researchers to curate code examples,... ...precise solutions, and refine AI-generated code across various...TrainingFor contractorsFlexible hours- ...Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems.... ...human-in-the-loop and AI-augmented workflows,... ...Data Scientists and AI/ML Research Engineers to design and validate evaluation...TrainingFull time
$100k - $132k
...Dynamics Mission SystemsCanada, were not just engineering technology were shaping the future of... ...systems, acoustic signal processing, training and simulation environments, and data... ...Experience working with or familiarity with AI/ML models is preferred. Preferred...TrainingFull timeCurrently hiringFlexible hours$85k - $120k
...Dynamics Mission SystemsCanada, were not just engineering technology were shaping the future of... ...technical guidance and on-the-job training to support personnel on system functionality... ...working with or familiarity with AI/ML models is preferred. Preferred Qualifications...TrainingFull timeFor contractorsFlexible hours$136k - $204k
...and improvement. Responsibilities Engineering Expectations (Core Capabilities) Design and build end-to-end AI systems integrating ML models, LLMs, APIs, and enterprise data... ..., and relevant education and/or training. The targeted pay range listed reflects...TrainingFull timeWork at office$20 - $40 per hour
...transformative project aimed at enhancing AI capabilities. Your role will involve... ...documents that will directly influence the training and performance of next-generation AI systems... ...Practicing lawyer, legal researcher, compliance consultant, or legal professional...TrainingRemote jobHourly payFor contractors$127k - $203k
...responsible for developing advanced AI and machine learning... ...Role: • Design, build, evaluate, enhance, and monitor machine... ...cases. • Oversee feature engineering, model training, validation, packaging, production... ...Collaborate closely with AI/ML engineering and development...TrainingFull timeWork at officeWorldwide3 days per week$170 - $190 per hour
...Join a dynamic healthcare AI partner dedicated to building advanced... ..., and case notes for AI model training. Identify and validate medical... ...with data scientists and engineers to enhance AI understanding of medical context. Model Evaluation & Feedback: Evaluate AI-generated...Training$139k - $181.5k
...building the future of work by engineering the world’s first truly converged, AI-native workspace that... ...streaming pipeline, and LLM feature-store... ...cloud cost management. AI/ML Infrastructure Customization... ...stores, embedding pipelines, training data arrays, and live model...TrainingPermanent employmentFull timeTemporary workWork at officeRemote workFlexible hoursShift work$100 - $150 per hour
...technology. You will play a crucial role in shaping the training of next-generation AI systems by providing high-quality, real-world legal insights... ...Apply legal reasoning and litigation experience to evaluate, refine, and enhance AI-generated legal content, including...TrainingRemote jobHourly payFor contractors$100k - $115k
...SystemsCanada, were not just engineering technology were shaping the... ...experimentation, and prototype evaluation for next-generation sensing... ...working with or familiarity with AI/ML models is preferred. It is... ..., and continuous training, youll have the resources you...Full timeFlexible hours$80 - $150 per hour
...impact project that shapes the future of AI systems by applying your physics... ...), you will play a critical role in evaluating and enhancing the training of next-generation AI models,... ...theoretical arguments generated by researchers or AI platforms. Detect errors, unjustified...TrainingRemote jobHourly payFor contractors$127k - $203k
...with our Foundry Research and Development team... ...through feature engineering, hyperparameter tuning... ...teams, including AI/ML engineering,... ..., and performance evaluation. • Experience working... ..., and exposure to LLM fine tuning and... ...mandatory security trainings in accordance with...Full timeWorldwide$146.18k - $219.27k
...About the role D-Wave is seeking a Staff Machine Learning Research Developer to work alongside our researchers, solutions architects, and software developers specializing in various domains (e.g., combinatorial optimization, graph theory, and quantum physics...Full timeLocal area$50 - $90 per hour
...will leverage your expertise to help train next-generation AI systems, shaping how these models learn... ...Establish benchmark guidelines to evaluate risk, support vulnerable individuals,... ...mental-health platforms, or relevant research is a plus. ~ Commitment to upholding...TrainingRemote jobHourly payFor contractors$40 - $50 per hour
...Join a dynamic team as an Engineering & Technical Documentation Specialist, where your expertise... ...understanding for next-generation AI systems. This remote position allows you... ...comprehensive rubrics and structured guidelines for evaluating visual and technical documentation...TrainingHourly payFor contractorsRemote work$40 - $50 per hour
...play a crucial role in shaping the training of next-generation AI systems. This position allows you to... ...reports, and operational procedures. Evaluate electrical troubleshooting scenarios... .... Backgrounds such as Electrical Engineer, Controls Engineer, Automation...TrainingRemote jobHourly payFor contractors$20 - $30 per hour
...to provide high-quality insights that contribute to the training of next-generation AI systems, shaping how these models learn and perform. Key... ...accuracy, and proper application of statutory rules. Evaluate individual tax returns (Form 1040) in light of current IRS...TrainingHourly payFor contractorsRemote work$50 - $90 per hour
...shaping the development of next-generation AI systems by providing high-quality, real-... ...create a robust mental health safety evaluation framework for vulnerable youth online, leveraging... ...and guidelines are grounded in current research and clinical best practices. Offer...TrainingRemote jobHourly payFor contractors$50 - $60 per hour
...will leverage your expertise to help train next-generation AI systems, shaping how models learn, reason... ...planning, sales analysis, market research, or growth enablement, who are eager... ...Key Responsibilities Review and evaluate marketing and sales strategy decks, campaign...TrainingRemote jobHourly payFor contractors$91k - $140k
...Artificial Intelligence (AI) and Machine Learning (ML) models that power... ...for the research and development of... ...extraction and feature engineering to validation, deployment... ..., and performance evaluation • Support model... ...mandatory security trainings in accordance with...Full timeWorldwide$111k - $160k
...Artificial Intelligence (AI) and Machine Learning (ML) models that... ...team also leads the research and development of... ...acquisition and feature engineering through... ...and machine learning evaluation • Ability to identify... ...mandatory security trainings in accordance with...Full timeWorldwide- ...Robots & Pencils is an applied AI engineering firm building the next frontier of business architecture... ...of data pipelines for production AI/ML systems, including embeddings, vector... ...data preparation, feature stores, and training/inference data flows Integrate data services...TrainingFull time
$190k
...something truly special. Staff Design Engineer As a Staff Design Engineer , you are... ...designers ship in code and integrate AI tools into their workflows—improving speed... ...budget for courses, certifications, and training to support your career growth. ~ Time off...TrainingRemote jobFull timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI/ML Research Engineer, LLM Training & Evaluation. Be the first to apply!
- ai prompt engineer Canada
- ai engineer Canada
- ai developer Canada
- machine learning remote Canada
- artificial intelligence - machine learning intern Canada
- data engineer machine learning Canada
- machine learning Canada
- research and development analyst Canada
- research evaluation Canada
- research and development internship Canada






