ML Engineer - LLM Evaluation & Automation
Grid Dynamics Holdings
Role Description
We are seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will design and build systems that automatically measure and improve the accuracy, relevance, and consistency of model outputs. You will lead initiatives to create evaluation pipelines, develop metrics, and deliver actionable insights for continuous improvements. This position requires strong technical expertise, analytical problem-solving abilities, and the capacity to manage projects across multiple cross-functional teams.
- Design and implement automated systems and pipelines for evaluating LLM outputs.
- Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations.
- Collaborate with Engineering teams to create automated logic checks and validation tools.
- Partner with Data Scientists to analyze evaluation results and optimize prompt and task structures.
- Provide feedback loops to ensure evaluation guidelines align with LLM-based assessments.
- Investigate how LLM-derived evaluations can enhance product reliability and user experience.
- Recommend refinements to prompt engineering, evaluation strategies, and automation tools.
- Stay informed on emerging trends in LLM evaluation, automated quality assessment, and AI toolchains.
- Continuously improve and expand automated evaluation processes based on industry best practices.
Qualifications
- 5+ years of experience in ML engineering, NLP, or AI/ML automation.
- Hands-on experience in prompt engineering and designing LLM-based evaluation systems is preferred.
- Strong understanding of machine learning principles with focus on NLP and advanced LLM capabilities (e.g., Chain-of-Thought, agentic workflows).
- Expertise in building automated evaluation or QA pipelines.
- Excellent analytical and problem-solving skills with experience in root cause and error pattern analysis.
- Proven project management and cross-functional collaboration experience.
- Excellent communication skills to convey complex insights to technical and non-technical audiences.
- Detail-oriented mindset with a focus on evaluation metrics, prompt design, and automation.
- Ability to quickly adapt to new business rules and evaluation guidelines across diverse product domains.
- Strong programming skills in Python and SQL.
- Experience with big data technologies like PySpark for data aggregation and sampling is a strong plus.
- Bachelor's/Master's degree in Computer Science/Engineering or a related field.
Benefits
- Opportunity to work on cutting-edge projects.
- Work with a highly motivated and dedicated team.
- Competitive salary.
- Flexible schedule.
- Benefits package - medical insurance, vision, dental, etc.
- Corporate social events.
- Professional development opportunities.
- Well-equipped office.
- ...Machine Learning Engineer (Llama AI Platform) Location: Remote (... ...custom AI agents, intelligent automation, and connected business systems... .... Fine-tune and evaluate LLM performance for business use... ...tuning open-source LLMs. ML Engineering and MLOps practices...SuggestedFull timeRemote work
$213k - $263k
...modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and you will... ...large foundation models to smaller real-time models. Explore LLM/VLM distillation recipes to maximally transfer the quality gains...SuggestedFull timeRemote work$238k - $302k
...inference, hierarchical learning, and robust evaluation. This role follows a hybrid work... ...you will report to a Senior Staff Software Engineer. You will: Work with a creative... ...experience ~5+ years of experience in ML engineering and applied Deep Learning, with...SuggestedFull timeRemote work$281k - $356k
...improve and speed up the evaluation and onboard developer... ...researchers and software engineers who are passionate... ...models and Generative AI (LLM/VLM) solutions. These solutions... ...triaging, introduce automation for high-volume... ...in Python and standard ML frameworks (e.g., JAX,...SuggestedFull time- ...expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build... ...Experimentation Experience implementing automated evaluation pipelines and test harnesses Experience...SuggestedFull time
- ...We are hiring a Founding AI/ML Engineer to architect and ship production... ...world’s largest and least automated industries — construction.... ...embeddings, reranking, chunking) Evaluation systems Orchestration... ...to real customers Hands-on LLM and/or computer vision experience...Remote jobFull timeH1bVisa sponsorship
$150k - $230k
...for a hands-on Machine Learning Engineer to drive the post-training... ...stability. Build and maintain evaluation and reward/verifier pipelines... ...Requirements Hands-on LLM post-training experience. You... ...Strong data engineering for ML. You can independently design...Full timeLocal areaWork from home$204k - $259k
...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently... ...models. Develop and rigorously evaluate metrics and methodologies for measuring the... ...focus on large-scale model development (LLM, VLM, or similar foundation models). ~...Full timeRemote work- ...your unique tastes, our Taste AI engine sifts through the noise to... ...Learning Engineer reporting to the LLM Research Lead, you will... ...includes building production-ready ML systems, experimenting with... ...explainability Experiment with and evaluate modern ML approaches (...Full timeRemote workFlexible hours
- ...quickly find answers and automate tasks. Powered by the... ...with Moveworks’ Reasoning Engine and natural language capabilities... ...help build cutting edge ML infrastructure for building and serving LLM’s at Moveworks. This... ...models(LLM), model evaluation and monitoring framework...Full timeWork at officeRemote workFlexible hours
$85 per hour
...Role Overview Evaluate and improve frontier AI coding models by completing... ...realistic machine learning engineering workflows, model training and inference systems, MLOps, and LLM application scenarios. Key... ...complete and evaluate complex ML and AI engineering tasks....Hourly payRemote work- ...Service to Marketing we rely on ML to ensure that guests and... ...and tools including LLM fine-tuning, alignment and... ...optimization, RAG/Search, LLM evaluation and testing automation, feedback-based learning and... ...principal machine learning engineer, you will be responsible for...Remote jobFull timeCasual workLive inWork at office
- ...to improve their business. Founded by engineers — and customer obsessed — we leap at every... ...Databricks, to bring enterprise-grade ML and AI personalization to every company... .... The impact you'll have: Evaluate ML and LLM approaches for CustomerLake's personalization...Full timeWork at office
- ...Webflow. We’re looking for an Engineer to join the ML Platform team at Synthesia.... ...making these systems more automation-friendly and agent-... ...that support model training, evaluation, and production serving.... ...Experience building agentic or LLM-powered internal tools....Remote work
$264k - $330k
...assistance to power real automation and decision-making.... ...Principal Machine Learning Engineer to help define and lead... ...for end-to-end ML systems: data collection, model training, evaluation, deployment, and inference... ...RL-optimize open source LLM and SLM for the RPM domain...Full timeRemote workFlexible hours$183k - $246k
...construction through intelligent automation. Despite being a $13+... ...100+ employees (50+ engineers), we’re scaling fast... ...Establish and improve evaluation and experimentation... ...and shipping production ML/AI systems. ~ Proven... ...with agentic systems, LLM-based applications, or...Work at officeRemote workFlexible hours$400k
...advanced math, and quantitative reasoning. As a core engineering team member, you will design scalable LLM-based services, implement model integrations, and translate... .... Expert-level proficiency in Python for AI/ML engineering. Hands-on experience deploying, fine-tuning...Full timeRemote work$199.5k - $276k
...The Team Upstart’s Applied LLM team is building foundational... ...generative AI for every product and engineering team across the company. This... ...is to bring the power of ML, particularly large language models... ...for prompt design, model evaluation, and experimentation across the...Full timeSummer workCurrently hiringLocal areaRemote workWork from home- .... Responsibilities: Own evaluation pipelines — design, build, and automate offline and live evals that keep... ...scale training and inference for LLM-class workloads; chase latency,... ...level PyTorch. Proven software engineer who loves ML; comfortable writing production...Full timeContract workFlexible hoursShift work
$144.7k - $261.3k
...Job Description The Senior ML Validation Research Engineer will lead applied machine learning research focused on... ...systems. This role centers on simulation-based evaluation, uncertainty modeling, scenario coverage automation, and transforming advanced ML research into...Local areaWork from homeFlexible hours$145.7k - $174.8k
...creating the future of financial automation so businesses can spend more... ...Join BILL's AI Product Engineering team and help shape the future... ...business requirements into scalable ML solutions Analyze complex... ...drive product innovation Evaluate, optimize, and monitor model...Full timeTemporary workWork at officeRemote workVisa sponsorshipFlexible hours$180k - $225k
...of intelligent technologies powered by automation, AI, ML, and knowledge graphs is accelerating.... ...are looking for a Machine Learning Engineer who will play a critical role in fine... ...frameworks that allow us to iterate and evaluate model versions (ranking, accuracy,...Remote jobFull timeFor contractorsWork experience placement$213k - $263k
...create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How do you mathematically... ...world is "real"? We are seeking visionary machine learning engineers and researchers to architect the scalable deep learning...Full timeRemote work$500 per month
...About Adelphi: Adelphi builds AI/ML-enabled secure data access and... ...eliminate data silos, build trust in automation without compromising security,... ...Forward Deployed Learning Engineer position requires a mix of software development, LLM Ops, and SecDevOps practices, resulting...Full timeRemote work$110k - $180k
...private market access, and smart automation to help you grow and protect... ...) Build systems for LLM orchestration , prompt management... ...workflow execution Develop evaluation frameworks for agent quality... ...ready systems Partner with ML and product teams to integrate...Full timeWork at officeRemote workRelocation- ...9 Principal Machine Learning Engineer HubSpot is an all-in-one marketing... ...revenue. From marketing automation to CRM, HubSpot offers a... ...Group at HubSpot delivers the ML and AI foundations that enable... ...opportunities through model development, evaluation, productionization,...Full timeWork at officeRemote work
$204k - $259k
...states. The Driver Understanding and Evaluation (DUE) team at Waymo is developing rich metrics... ...are looking for researchers and software engineers who are passionate about developing... ...Architect and implement scalable and robust ML pipelines for training, evaluating, and...Full time$140k - $230k
...software platform, safety-first automated driving technology and Toyota... ...the-art of machine learning (ML) for perception, prediction,... ...other software and hardware engineers and researchers to tackle... ...training, ablation studies, evaluation, deployment, inference optimization...Full timeTemporary workFlexible hours$170k - $216k
...insight tools, improve and speed up the evaluation and onboard developer journeys. It will combine... ...are looking for researchers and software engineers who are passionate about developing... ...scale large distributed systems covering the ML lifecycle, supporting planet-scale...Full time$266k - $372.4k
...quickly build a new kind of search engine. As a Senior Staff, you will... ..., features, measurement, LLM-based answers, RAG etc, partnering... .... You will also deploy ML models, integrate LLMs, and ensure... ...LLM in production, including evaluation, tuning and deployment. ~ In...Full timeFor contractorsWork experience placementRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer - LLM Evaluation & Automation. Be the first to apply!
- computer vision machine learning engineer Remote
- junior machine learning research engineer Remote
- machine learning software engineer Remote
- machine learning ai engineer Remote
- data scientist machine learning engineer Remote
- machine learning engineer Remote
- graduate machine learning engineer Remote
- ai ml engineer Remote
- senior ml engineer Remote
- infrastructure automation engineer Remote




