Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer - LLM Evaluation & Automation

Full-time

Grid Dynamics Holdings

Role Description

We are seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will design and build systems that automatically measure and improve the accuracy, relevance, and consistency of model outputs. You will lead initiatives to create evaluation pipelines, develop metrics, and deliver actionable insights for continuous improvements. This position requires strong technical expertise, analytical problem-solving abilities, and the capacity to manage projects across multiple cross-functional teams.

  • Design and implement automated systems and pipelines for evaluating LLM outputs.
  • Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations.
  • Collaborate with Engineering teams to create automated logic checks and validation tools.
  • Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures.
  • Provide feedback loops to ensure evaluation guidelines align with LLM-based assessments.
  • Investigate how LLM-derived evaluations can enhance product reliability and user experience.
  • Recommend refinements to prompt engineering, evaluation strategies, and automation tools.
  • Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains.
  • Continuously improve and expand automated evaluation processes based on industry best practices.

Qualifications

  • 5+ years of experience in ML engineering, NLP, or AI/ML automation.
  • Hands-on experience in prompt engineering and designing LLM-based evaluation systems is preferred.
  • Strong understanding of machine learning principles with focus on NLP and advanced LLM capabilities (e.g., Chain-of-Thought, agentic workflows).
  • Expertise in building automated evaluation or QA pipelines.
  • Excellent analytical and problem-solving skills with experience in root cause and error pattern analysis.
  • Proven project management and cross-functional collaboration experience.
  • Excellent communication skills to convey complex insights to technical and non-technical audiences.
  • Detail-oriented mindset with a focus on evaluation metrics, prompt design, and automation.
  • Ability to quickly adapt to new business rules and evaluation guidelines across diverse product domains.
  • Strong programming skills in Python and SQL.
  • Experience with big data technologies like PySpark for data aggregation and sampling is a strong plus.
  • Bachelor's/Master's degree in Computer Science/Engineering or a related field.

Benefits

  • Opportunity to work on cutting-edge projects.
  • Work with a highly motivated and dedicated team.
  • Competitive salary.
  • Flexible schedule.
  • Benefits package - medical insurance, vision, dental, etc.
  • Corporate social events.
  • Professional development opportunities.
  • Well-equipped office.
Vacancy posted 20 days ago
Similar jobs that could be interesting for youBased on the ML Engineer - LLM Evaluation & Automation in Remote vacancy
  •  ...Machine Learning Engineer (Llama AI Platform) Location: Remote (...  ...custom AI agents, intelligent automation, and connected business systems...  .... Fine-tune and evaluate LLM performance for business use...  ...tuning open-source LLMs. ML Engineering and MLOps practices... 
    Suggested
    Full time
    Remote work

    Performacentric

    Denver, CO
    1 day ago
  • $213k - $263k

     ...modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and you will...  ...large foundation models to smaller real-time models. Explore LLM/VLM distillation recipes to maximally transfer the quality gains... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  • $238k - $302k

     ...inference, hierarchical learning, and robust evaluation. This role follows a hybrid work...  ...you will report to a Senior Staff Software Engineer.   You will: Work with a creative...  ...experience ~5+ years of experience in ML engineering and applied Deep Learning, with... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  • $281k - $356k

     ...improve and speed up the evaluation and onboard developer...  ...researchers and software engineers who are passionate...  ...models and Generative AI (LLM/VLM) solutions. These solutions...  ...triaging, introduce automation for high-volume...  ...in Python and standard ML frameworks (e.g., JAX,... 
    Suggested
    Full time

    Waymo

    Remote
    1 day ago
  •  ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build...  ...Experimentation Experience implementing automated evaluation pipelines and test harnesses Experience... 
    Suggested
    Full time

    Innodata

    Remote
    10 days ago
  •  ...We are hiring a Founding AI/ML Engineer  to architect and ship production...  ...world’s largest and least automated industries — construction....  ...embeddings, reranking, chunking) Evaluation systems Orchestration...  ...to real customers Hands-on LLM and/or computer vision experience... 
    Remote job
    Full time
    H1b
    Visa sponsorship

    Pulserise Technologies

    United States
    1 day ago
  • $150k - $230k

     ...for a hands-on Machine Learning Engineer to drive the post-training...  ...stability. Build and maintain evaluation and reward/verifier pipelines...  ...Requirements Hands-on LLM post-training experience. You...  ...Strong data engineering for ML. You can independently design... 
    Full time
    Local area
    Work from home

    News Break

    Remote
    1 day ago
  • $204k - $259k

     ...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently...  ...models. Develop and rigorously evaluate metrics and methodologies for measuring the...  ...focus on large-scale model development (LLM, VLM, or similar foundation models). ~... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...your unique tastes, our Taste AI engine sifts through the noise to...  ...Learning Engineer reporting to the LLM Research Lead, you will...  ...includes building production-ready ML systems, experimenting with...  ...explainability Experiment with and evaluate modern ML approaches (... 
    Full time
    Remote work
    Flexible hours

    Qloo

    New York, NY
    1 day ago
  •  ...quickly find answers and automate tasks. Powered by the...  ...with Moveworks’ Reasoning Engine and natural language capabilities...  ...help build cutting edge ML infrastructure for building and serving LLM’s at Moveworks. This...  ...models(LLM), model evaluation and monitoring framework... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Servicenow

    Remote
    1 day ago
  • $85 per hour

     ...Role Overview Evaluate and improve frontier AI coding models by completing...  ...realistic machine learning engineering workflows, model training and inference systems, MLOps, and LLM application scenarios. Key...  ...complete and evaluate complex ML and AI engineering tasks.... 
    Hourly pay
    Remote work

    SaidGig

    United States
    5 days ago
  •  ...Service to Marketing we rely on ML to ensure that guests and...  ...and tools including LLM fine-tuning, alignment and...  ...optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning and...  ...principal machine learning engineer, you will be responsible for... 
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    United States
    1 day ago
  •  ...to improve their business. Founded by engineers — and customer obsessed — we leap at every...  ...Databricks, to bring enterprise-grade ML and AI personalization to every company...  ....   The impact you'll have: Evaluate ML and LLM approaches for CustomerLake's personalization... 
    Full time
    Work at office

    Databricks

    Remote
    1 day ago
  •  ...Webflow. We’re looking for an Engineer to join the ML Platform team at Synthesia....  ...making these systems more automation-friendly and agent-...  ...that support model training, evaluation, and production serving....  ...Experience building agentic or LLM-powered internal tools.... 
    Remote work

    Synthesia

    United States
    3 days ago
  • $264k - $330k

     ...assistance to power real automation and decision-making....  ...Principal Machine Learning Engineer to help define and lead...  ...for end-to-end ML systems: data collection, model training, evaluation, deployment, and inference...  ...RL-optimize open source LLM and SLM for the RPM domain... 
    Full time
    Remote work
    Flexible hours

    AppFolio

    United States
    3 days ago
  • $183k - $246k

     ...construction through intelligent automation. Despite being a $13+...  ...100+ employees (50+ engineers), we’re scaling fast...  ...Establish and improve evaluation and experimentation...  ...and shipping production ML/AI systems. ~ Proven...  ...with agentic systems, LLM-based applications, or... 
    Work at office
    Remote work
    Flexible hours

    Trunk Tools, Inc.

    New York, NY
    12 days ago
  • $400k

     ...advanced math, and quantitative reasoning. As a core engineering team member, you will design scalable LLM-based services, implement model integrations, and translate...  .... Expert-level proficiency in Python for AI/ML engineering. Hands-on experience deploying, fine-tuning... 
    Full time
    Remote work

    SaidGig

    United States
    3 days ago
  • $199.5k - $276k

     ...The Team Upstart’s Applied LLM team is building foundational...  ...generative AI for every product and engineering team across the company. This...  ...is to bring the power of ML, particularly large language models...  ...for prompt design, model evaluation, and experimentation across the... 
    Full time
    Summer work
    Currently hiring
    Local area
    Remote work
    Work from home

    Upstart

    United States
    5 days ago
  •  .... Responsibilities: Own evaluation pipelines — design, build, and automate offline and live evals that keep...  ...scale training and inference for LLM-class workloads; chase latency,...  ...level PyTorch. Proven software engineer who loves ML; comfortable writing production... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    1 day ago
  • $144.7k - $261.3k

     ...Job Description The Senior ML Validation Research Engineer will lead applied machine learning research focused on...  ...systems. This role centers on simulation-based evaluation, uncertainty modeling, scenario coverage automation, and transforming advanced ML research into... 
    Local area
    Work from home
    Flexible hours

    General Motors

    Helena, MT
    1 day ago
  • $145.7k - $174.8k

     ...creating the future of financial automation so businesses can spend more...  ...Join BILL's AI Product Engineering team and help shape the future...  ...business requirements into scalable ML solutions Analyze complex...  ...drive product innovation Evaluate, optimize, and monitor model... 
    Full time
    Temporary work
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours

    Bill

    Draper, UT
    1 day ago
  • $180k - $225k

     ...of intelligent technologies powered by automation, AI, ML, and knowledge graphs is accelerating....  ...are looking for a Machine Learning Engineer who will play a critical role in fine...  ...frameworks that allow us to iterate and evaluate model versions (ranking, accuracy,... 
    Remote job
    Full time
    For contractors
    Work experience placement

    Adapter

    Remote
    1 day ago
  • $213k - $263k

     ...create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How do you mathematically...  ...world is "real"? We are seeking visionary machine learning engineers and researchers to architect the scalable deep learning... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $500 per month

     ...About Adelphi: Adelphi builds AI/ML-enabled secure data access and...  ...eliminate data silos, build trust in automation without compromising security,...  ...Forward Deployed Learning Engineer position requires a mix of software development, LLM Ops, and SecDevOps practices, resulting... 
    Full time
    Remote work

    Adelphi Inc

    Washington DC
    1 day ago
  • $110k - $180k

     ...private market access, and smart automation to help you grow and protect...  ...) Build systems for LLM orchestration , prompt management...  ...workflow execution Develop evaluation frameworks for agent quality...  ...ready systems Partner with ML and product teams to integrate... 
    Full time
    Work at office
    Remote work
    Relocation

    Arta Finance

    Bay, AR
    1 day ago
  •  ...9 Principal Machine Learning Engineer HubSpot is an all-in-one marketing...  ...revenue. From marketing automation to CRM, HubSpot offers a...  ...Group at HubSpot delivers the ML and AI foundations that enable...  ...opportunities through model development, evaluation, productionization,... 
    Full time
    Work at office
    Remote work

    Hubspot

    Remote
    1 day ago
  • $204k - $259k

     ...states. The Driver Understanding and Evaluation (DUE) team at Waymo is developing rich metrics...  ...are looking for researchers and software engineers who are passionate about developing...  ...Architect and implement scalable and robust ML pipelines for training, evaluating, and... 
    Full time

    Waymo

    Remote
    1 day ago
  • $140k - $230k

     ...software platform, safety-first automated driving technology and Toyota...  ...the-art of machine learning (ML) for perception, prediction,...  ...other software and hardware engineers and researchers to tackle...  ...training, ablation studies, evaluation, deployment, inference optimization... 
    Full time
    Temporary work
    Flexible hours

    Woven By Toyota

    Remote
    1 day ago
  • $170k - $216k

     ...insight tools, improve and speed up the evaluation and onboard developer journeys. It will combine...  ...are looking for researchers and software engineers who are passionate about developing...  ...scale large distributed systems covering the ML lifecycle, supporting planet-scale... 
    Full time

    Waymo

    Remote
    1 day ago
  • $266k - $372.4k

     ...quickly build a new kind of search engine. As a Senior Staff, you will...  ..., features, measurement, LLM-based answers, RAG etc, partnering...  .... You will also deploy ML models, integrate LLMs, and ensure...  ...LLM in production, including evaluation, tuning and deployment. ~ In... 
    Full time
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    United States
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer - LLM Evaluation & Automation. Be the first to apply!