Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer - LLM Evaluation & Automation

Grid Dynamics Holdings

We are seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will design and build systems that automatically measure and improve the accuracy, relevance, and consistency of model outputs. You will lead initiatives to create evaluation pipelines, develop metrics, and deliver actionable insights for continuous improvements. This position requires strong technical expertise, analytical problem-solving abilities, and the capacity to manage projects across multiple cross-functional teams.

Essential functions


  • Design and implement automated systems and pipelines for evaluating LLM outputs.
  • Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations
  • Collaborate with Engineering teams to create automated logic checks and validation tools.
  • Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures.
  • Provide feedback loops to ensure evaluation guidelines align with LLM-based assessments.
  • Investigate how LLM-derived evaluations can enhance product reliability and user experience.
  • Recommend refinements to prompt engineering, evaluation strategies, and automation tools.
  • Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains.
  • Continuously improve and expand automated evaluation processes based on industry best practices.
Qualifications
  • 5+ years of experience in ML engineering, NLP, or AI/ML automation.
  • Hands-on experience in prompt engineering and designing LLM-based evaluation systems is preferred
  • Strong understanding of machine learning principles with focus on NLP and advanced LLM capabilities (e.g., Chain-of-Thought, agentic workflows)
  • Expertise in building automated evaluation or QA pipelines.
  • Excellent analytical and problem-solving skills with experience in root cause and error pattern analysis.
  • Proven project management and cross-functional collaboration experience.
  • Excellent communication skills to convey complex insights to technical and non-technical audiences.
  • Detail-oriented mindset with a focus on evaluation metrics, prompt design, and automation.
  • Ability to quickly adapt to new business rules and evaluation guidelines across diverse product domains.
  • Strong programming skills in Python and SQL.
  • Experience with big data technologies like PySpark for data aggregation and sampling is a strong plus
  • Bachelor's/Master's degree in Computer Science/ Engineering or a related field.
We offer
  • Opportunity to work on cutting-edge projects
  • Work with a highly motivated and dedicated team
  • Competitive salary
  • Flexible schedule
  • Benefits package - medical insurance, vision, dental, etc.
  • Corporate social events
  • Professional development opportunities
  • Well-equipped office

About us
Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI , supported by profound expertise and ongoing investment in data , analytics , cloud & DevOps , application modernization and customer experience . Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the ML Engineer - LLM Evaluation & Automation in United States vacancy
  •  ...and real-world responsibility in mind. Our ML team comes from a culture of academic...  ...frontier research for their next generation of LLM products. Join us if you: Wish to work...  ...ML advancement. Responsibilities Own LLM evaluation processes and methods with a focus on generating... 
    Suggested
    Local area
    Shift work

    Capitolis

    San Francisco, CA
    23 hours ago
  •  ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build...  ...Experimentation Experience implementing automated evaluation pipelines and test harnesses Experience... 
    Suggested
    Full time

    Innodata

    Remote
    10 days ago
  •  ...Machine Learning Engineer (Llama AI Platform) Location: Remote (...  ...custom AI agents, intelligent automation, and connected business systems...  .... Fine-tune and evaluate LLM performance for business use...  ...tuning open-source LLMs. ML Engineering and MLOps practices... 
    Suggested
    Full time
    Remote work

    Performacentric

    Denver, CO
    9 hours ago
  • $213k - $263k

     ...modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and you will...  ...large foundation models to smaller real-time models. Explore LLM/VLM distillation recipes to maximally transfer the quality gains... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    9 hours ago
  • $238k - $302k

     ...inference, hierarchical learning, and robust evaluation. This role follows a hybrid work...  ...you will report to a Senior Staff Software Engineer.   You will: Work with a creative...  ...experience ~5+ years of experience in ML engineering and applied Deep Learning, with... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    9 hours ago
  • $281k - $356k

     ...improve and speed up the evaluation and onboard developer...  ...researchers and software engineers who are passionate...  ...models and Generative AI (LLM/VLM) solutions. These solutions...  ...triaging, introduce automation for high-volume...  ...in Python and standard ML frameworks (e.g., JAX,... 
    Full time

    Waymo

    Remote
    9 hours ago
  • $500 per month

     ...Deployed Senior Machine Learning Engineer Adelphi builds AI/ML-enabled secure data access...  ...silos, build trust in automation without compromising...  ...of software development, LLM Ops, and secdevops practices...  ...AutoGen, or similar) and agent evaluation / observability tooling.... 
    Remote work

    Adelphi

    United States
    5 days ago
  • $183k - $246k

     ...construction through intelligent automation. Despite being a $13+...  ...100+ employees (50+ engineers), we’re scaling fast...  ...Establish and improve evaluation and experimentation...  ...and shipping production ML/AI systems. ~ Proven...  ...with agentic systems, LLM-based applications, or... 
    Work at office
    Remote work
    Flexible hours

    Trunk Tools, Inc.

    New York, NY
    more than 2 months ago
  •  ....As a Senior Lead Software Engineer - Python/AWS/AI/LLM at JPMorganChase within the...  ...CycleCollaborate with firmwide AI/ML teams, Business and Product...  ...coding, peer review, automated testing) and promoting...  ...workflow toolset such as tracing, evaluations, and guardrails; Must have... 
    For contractors

    JP Morgan Chase

    Jersey City, NJ
    1 day ago
  • General Information Job Title ML Staff Engineer - LLM & Production Systems Job ID 107242 Work Areas Technology & Engineering...  ...robust data contracts and monitoring practicesBuild evaluation frameworks including dashboards, regression testing, and slice... 
    Permanent employment
    Full time
    Work at office
    Local area
    1 day per week

    Bain & Company

    Dallas, TX
    2 days ago
  • ML Engineer | Nox Metals | Detroit, MI American factories deserve a supply...  .... We use software and automation to supply metal to American factories...  ...at what price Build NLP and LLM features for sales order...  ...collection, labeling, training, evaluation, deployment, monitoring,... 
    Full time
    Immediate start
    Shift work

    Nox Metals

    Detroit, MI
    3 days ago
  •  ...science, and software engineering to address high-impact...  ...Machine Learning Engineer (Evaluation & Experimentation) at...  ...experiments and evaluate LLM outputs. This role is...  ...outputs at scale Build automated evaluation workflows...  ...functionally Work closely with ML systems engineers to... 

    Cynnovative

    Arlington, VA
    3 days ago
  •  ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build...  ...Experimentation Experience implementing automated evaluation pipelines and test harnesses Experience... 
    Full time

    Innodata

    Remote
    10 days ago
  •  ...We are hiring a Founding AI/ML Engineer  to architect and ship production...  ...world’s largest and least automated industries — construction....  ...embeddings, reranking, chunking) Evaluation systems Orchestration...  ...to real customers Hands-on LLM and/or computer vision experience... 
    Remote job
    Full time
    H1b
    Visa sponsorship

    Pulserise Technologies

    United States
    9 hours ago
  • $148k - $184k

     ...help define how we measure and trust the LLM systems that power Paradigm's clinical...  ...at the intersection of applied AI evaluation, analytics engineering, and clinical research, partnering with...  ...a data scientist, analytics engineer, ML/data professional, or other highly analytical... 
    Full time

    Paradigm Health

    Remote
    9 hours ago
  • $204k - $259k

     ...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently...  ...models. Develop and rigorously evaluate metrics and methodologies for measuring the...  ...focus on large-scale model development (LLM, VLM, or similar foundation models). ~... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    9 hours ago
  •  ...quickly find answers and automate tasks. Powered by the...  ...with Moveworks’ Reasoning Engine and natural language capabilities...  ...help build cutting edge ML infrastructure for building and serving LLM’s at Moveworks. This...  ...models(LLM), model evaluation and monitoring framework... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Servicenow

    Remote
    9 hours ago
  •  ...your unique tastes, our Taste AI engine sifts through the noise to...  ...Learning Engineer reporting to the LLM Research Lead, you will...  ...includes building production-ready ML systems, experimenting with...  ...explainability Experiment with and evaluate modern ML approaches (... 
    Full time
    Remote work
    Flexible hours

    Qloo

    New York, NY
    9 hours ago
  • $120k - $185k

     ...proprietary software, automation, and advanced manufacturing...  ...for a Machine Learning Engineer to train custom models...  ..., cost estimation, evaluating the complexity and difficulty...  ...and iterate on custom ML models using Layup's...  ...than natural language or LLM-centric work  Bonus... 
    Permanent employment
    Full time
    Temporary work
    Local area

    Layup Parts

    Remote
    9 hours ago
  •  ...category-defining AI workflow automation platform that...  ...for a Machine Learning Engineer to design, build, and deploy production-grade ML systems that power the...  ...pipelines for training, evaluation, monitoring, and inference...  ...services using modern NLP, LLM, classification,... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    9 hours ago
  • $150k - $230k

     ...for a hands-on Machine Learning Engineer to drive the post-training...  ...stability. Build and maintain evaluation and reward/verifier pipelines...  ...Requirements Hands-on LLM post-training experience. You...  ...Strong data engineering for ML. You can independently design... 
    Full time
    Local area
    Work from home

    News Break

    Remote
    9 hours ago
  •  ...underwriting expertise, data, and automation work together to...  .... You will build the ML systems that carry us...  ...first Machine Learning Engineer, embedded in the Fully...  ...confidence scoring and evaluation frameworks that define...  ...frameworks or multi-step LLM orchestration (LangChain... 
    Full time
    Work at office

    Shepherd

    San Francisco, CA
    9 hours ago
  • $170k - $210k

     ...building a production-grade AI/ML capability spanning intelligent automation, agentic systems,...  ...predictive analytics, and data engineering. As AI/ML Engineering...  ...-the-loop controls, and evaluation frameworks Own the AI/ML...  ...AI systems using modern LLM orchestration frameworks... 
    For contractors
    Work experience placement

    Corsair

    Milpitas, CA
    3 days ago
  •  ...SoftwareClient: WiproContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: ML Engineer with LLMLocation: Sunnyvale, CA(onsite)Job Description:6-8 years of experience in machine learning and LLM, with a proven track record in image processing and analysis.Development... 

    SRI Tech

    Sunnyvale, CA
    1 day ago
  • $174.72k - $295.68k

     ...mission is to build strong foundation for LLM deployment and quality sign-off for next...  ...and prove model feasibility.Curate evaluation datasets and establish a comprehensive metric...  ....Strong Python programming and software engineering skills.Ability to work effectively... 
    Full time

    XPENG Motors

    Santa Clara, CA
    4 hours ago
  • Driverai is seeking an Applied Data Scientist with expertise in LLM evaluation to join its innovative team in Austin, TX. This role focuses...  ...quality metrics, create evaluation datasets, and establish automated quality signals for content generation pipelines. Driverai offers... 
    Remote job

    Driverai

    Austin, TX
    1 day ago
  • $175k - $275k

     ...location) - Senior - Product & Engineering - $175k - $275k Applied Data Scientist, LLM Evaluation Introduction At Driver, we’re building...  ...balance human judgment with automated signals. This role builds the...  ...— 5 years in applied science, ML engineering, or data science... 
    Remote job
    Full time
    Flexible hours

    Driverai

    Austin, TX
    1 day ago
  •  ...to improve their business. Founded by engineers — and customer obsessed — we leap at every...  ...Databricks, to bring enterprise-grade ML and AI personalization to every company...  ....   The impact you'll have: Evaluate ML and LLM approaches for CustomerLake's personalization... 
    Full time
    Work at office

    Databricks

    Remote
    9 hours ago
  •  ...Service to Marketing we rely on ML to ensure that guests and...  ...and tools including LLM fine-tuning, alignment and...  ...optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning and...  ...principal machine learning engineer, you will be responsible for... 
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    United States
    9 hours ago
  • $120.8k - $193.3k

    DescriptionJob Title: Sr. Engineer, AI/ML Apps Engineering Job Location:...  ...autonomous systems, industrial automation, drones, and edge AI...  ...development lifecycle—from initial evaluations and proof-of-concepts...  ..., multimodal AI systems, or LLM-based edge applications.Background... 
    Full time
    Work at office

    SiMa Technologies

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer - LLM Evaluation & Automation. Be the first to apply!