ML Engineer - LLM Evaluation & Automation
Grid Dynamics Holdings
We are seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will design and build systems that automatically measure and improve the accuracy, relevance, and consistency of model outputs. You will lead initiatives to create evaluation pipelines, develop metrics, and deliver actionable insights for continuous improvements. This position requires strong technical expertise, analytical problem-solving abilities, and the capacity to manage projects across multiple cross-functional teams.
Essential functions- Design and implement automated systems and pipelines for evaluating LLM outputs.
- Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations
- Collaborate with Engineering teams to create automated logic checks and validation tools.
- Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures.
- Provide feedback loops to ensure evaluation guidelines align with LLM-based assessments.
- Investigate how LLM-derived evaluations can enhance product reliability and user experience.
- Recommend refinements to prompt engineering, evaluation strategies, and automation tools.
- Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains.
- Continuously improve and expand automated evaluation processes based on industry best practices.
- 5+ years of experience in ML engineering, NLP, or AI/ML automation.
- Hands-on experience in prompt engineering and designing LLM-based evaluation systems is preferred
- Strong understanding of machine learning principles with focus on NLP and advanced LLM capabilities (e.g., Chain-of-Thought, agentic workflows)
- Expertise in building automated evaluation or QA pipelines.
- Excellent analytical and problem-solving skills with experience in root cause and error pattern analysis.
- Proven project management and cross-functional collaboration experience.
- Excellent communication skills to convey complex insights to technical and non-technical audiences.
- Detail-oriented mindset with a focus on evaluation metrics, prompt design, and automation.
- Ability to quickly adapt to new business rules and evaluation guidelines across diverse product domains.
- Strong programming skills in Python and SQL.
- Experience with big data technologies like PySpark for data aggregation and sampling is a strong plus
- Bachelor's/Master's degree in Computer Science/ Engineering or a related field.
- Opportunity to work on cutting-edge projects
- Work with a highly motivated and dedicated team
- Competitive salary
- Flexible schedule
- Benefits package - medical insurance, vision, dental, etc.
- Corporate social events
- Professional development opportunities
- Well-equipped office
About us
Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI , supported by profound expertise and ongoing investment in data , analytics , cloud & DevOps , application modernization and customer experience . Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.
- ...and real-world responsibility in mind. Our ML team comes from a culture of academic... ...frontier research for their next generation of LLM products. Join us if you: Wish to work... ...ML advancement. Responsibilities Own LLM evaluation processes and methods with a focus on generating...SuggestedLocal areaShift work
- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build... ...Experimentation Experience implementing automated evaluation pipelines and test harnesses Experience...SuggestedFull time
- ...Machine Learning Engineer (Llama AI Platform) Location: Remote (... ...custom AI agents, intelligent automation, and connected business systems... .... Fine-tune and evaluate LLM performance for business use... ...tuning open-source LLMs. ML Engineering and MLOps practices...SuggestedFull timeRemote work
$213k - $263k
...modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and you will... ...large foundation models to smaller real-time models. Explore LLM/VLM distillation recipes to maximally transfer the quality gains...SuggestedFull timeRemote work$238k - $302k
...inference, hierarchical learning, and robust evaluation. This role follows a hybrid work... ...you will report to a Senior Staff Software Engineer. You will: Work with a creative... ...experience ~5+ years of experience in ML engineering and applied Deep Learning, with...SuggestedFull timeRemote work$281k - $356k
...improve and speed up the evaluation and onboard developer... ...researchers and software engineers who are passionate... ...models and Generative AI (LLM/VLM) solutions. These solutions... ...triaging, introduce automation for high-volume... ...in Python and standard ML frameworks (e.g., JAX,...Full time$500 per month
...Deployed Senior Machine Learning Engineer Adelphi builds AI/ML-enabled secure data access... ...silos, build trust in automation without compromising... ...of software development, LLM Ops, and secdevops practices... ...AutoGen, or similar) and agent evaluation / observability tooling....Remote work$183k - $246k
...construction through intelligent automation. Despite being a $13+... ...100+ employees (50+ engineers), we’re scaling fast... ...Establish and improve evaluation and experimentation... ...and shipping production ML/AI systems. ~ Proven... ...with agentic systems, LLM-based applications, or...Work at officeRemote workFlexible hours- ....As a Senior Lead Software Engineer - Python/AWS/AI/LLM at JPMorganChase within the... ...CycleCollaborate with firmwide AI/ML teams, Business and Product... ...coding, peer review, automated testing) and promoting... ...workflow toolset such as tracing, evaluations, and guardrails; Must have...For contractors
- General Information Job Title ML Staff Engineer - LLM & Production Systems Job ID 107242 Work Areas Technology & Engineering... ...robust data contracts and monitoring practicesBuild evaluation frameworks including dashboards, regression testing, and slice...Permanent employmentFull timeWork at officeLocal area1 day per week
- ML Engineer | Nox Metals | Detroit, MI American factories deserve a supply... .... We use software and automation to supply metal to American factories... ...at what price Build NLP and LLM features for sales order... ...collection, labeling, training, evaluation, deployment, monitoring,...Full timeImmediate startShift work
- ...science, and software engineering to address high-impact... ...Machine Learning Engineer (Evaluation & Experimentation) at... ...experiments and evaluate LLM outputs. This role is... ...outputs at scale Build automated evaluation workflows... ...functionally Work closely with ML systems engineers to...
- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build... ...Experimentation Experience implementing automated evaluation pipelines and test harnesses Experience...Full time
- ...We are hiring a Founding AI/ML Engineer to architect and ship production... ...world’s largest and least automated industries — construction.... ...embeddings, reranking, chunking) Evaluation systems Orchestration... ...to real customers Hands-on LLM and/or computer vision experience...Remote jobFull timeH1bVisa sponsorship
$148k - $184k
...help define how we measure and trust the LLM systems that power Paradigm's clinical... ...at the intersection of applied AI evaluation, analytics engineering, and clinical research, partnering with... ...a data scientist, analytics engineer, ML/data professional, or other highly analytical...Full time$204k - $259k
...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently... ...models. Develop and rigorously evaluate metrics and methodologies for measuring the... ...focus on large-scale model development (LLM, VLM, or similar foundation models). ~...Full timeRemote work- ...quickly find answers and automate tasks. Powered by the... ...with Moveworks’ Reasoning Engine and natural language capabilities... ...help build cutting edge ML infrastructure for building and serving LLM’s at Moveworks. This... ...models(LLM), model evaluation and monitoring framework...Full timeWork at officeRemote workFlexible hours
- ...your unique tastes, our Taste AI engine sifts through the noise to... ...Learning Engineer reporting to the LLM Research Lead, you will... ...includes building production-ready ML systems, experimenting with... ...explainability Experiment with and evaluate modern ML approaches (...Full timeRemote workFlexible hours
$120k - $185k
...proprietary software, automation, and advanced manufacturing... ...for a Machine Learning Engineer to train custom models... ..., cost estimation, evaluating the complexity and difficulty... ...and iterate on custom ML models using Layup's... ...than natural language or LLM-centric work Bonus...Permanent employmentFull timeTemporary workLocal area- ...category-defining AI workflow automation platform that... ...for a Machine Learning Engineer to design, build, and deploy production-grade ML systems that power the... ...pipelines for training, evaluation, monitoring, and inference... ...services using modern NLP, LLM, classification,...Full timeWork at officeRemote workFlexible hours2 days per week
$150k - $230k
...for a hands-on Machine Learning Engineer to drive the post-training... ...stability. Build and maintain evaluation and reward/verifier pipelines... ...Requirements Hands-on LLM post-training experience. You... ...Strong data engineering for ML. You can independently design...Full timeLocal areaWork from home- ...underwriting expertise, data, and automation work together to... .... You will build the ML systems that carry us... ...first Machine Learning Engineer, embedded in the Fully... ...confidence scoring and evaluation frameworks that define... ...frameworks or multi-step LLM orchestration (LangChain...Full timeWork at office
$170k - $210k
...building a production-grade AI/ML capability spanning intelligent automation, agentic systems,... ...predictive analytics, and data engineering. As AI/ML Engineering... ...-the-loop controls, and evaluation frameworks Own the AI/ML... ...AI systems using modern LLM orchestration frameworks...For contractorsWork experience placement- ...SoftwareClient: WiproContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: ML Engineer with LLMLocation: Sunnyvale, CA(onsite)Job Description:6-8 years of experience in machine learning and LLM, with a proven track record in image processing and analysis.Development...
$174.72k - $295.68k
...mission is to build strong foundation for LLM deployment and quality sign-off for next... ...and prove model feasibility.Curate evaluation datasets and establish a comprehensive metric... ....Strong Python programming and software engineering skills.Ability to work effectively...Full time- Driverai is seeking an Applied Data Scientist with expertise in LLM evaluation to join its innovative team in Austin, TX. This role focuses... ...quality metrics, create evaluation datasets, and establish automated quality signals for content generation pipelines. Driverai offers...Remote job
$175k - $275k
...location) - Senior - Product & Engineering - $175k - $275k Applied Data Scientist, LLM Evaluation Introduction At Driver, we’re building... ...balance human judgment with automated signals. This role builds the... ...— 5 years in applied science, ML engineering, or data science...Remote jobFull timeFlexible hours- ...to improve their business. Founded by engineers — and customer obsessed — we leap at every... ...Databricks, to bring enterprise-grade ML and AI personalization to every company... .... The impact you'll have: Evaluate ML and LLM approaches for CustomerLake's personalization...Full timeWork at office
- ...Service to Marketing we rely on ML to ensure that guests and... ...and tools including LLM fine-tuning, alignment and... ...optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning and... ...principal machine learning engineer, you will be responsible for...Remote jobFull timeCasual workLive inWork at office
$120.8k - $193.3k
DescriptionJob Title: Sr. Engineer, AI/ML Apps Engineering Job Location:... ...autonomous systems, industrial automation, drones, and edge AI... ...development lifecycle—from initial evaluations and proof-of-concepts... ..., multimodal AI systems, or LLM-based edge applications.Background...Full timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer - LLM Evaluation & Automation. Be the first to apply!
- junior machine learning engineer United States
- graduate machine learning engineer United States
- senior ml engineer United States
- data scientist machine learning engineer United States
- machine learning engineer United States
- lead machine learning engineer United States
- machine learning ai engineer United States
- entry level machine learning engineer United States
- machine learning software engineer United States
- ai ml engineer United States





