AI Evaluations Engineer
VeeRteq Solutions Inc.
Role: AI Evaluations Engineer (Senior)
Location: US (remote)
Duration: 12 Months +
Key Responsibilities
- Design and operate end-to-end evaluation frameworks for agentic AI systems (offline regression, online monitoring, A/B testing)
- Define quality rubrics, scoring architectures, and evaluation standards across multi-agent workflows
- Build LLM-as-judge and Agent-as-judge pipelines assessing trajectory, tool-call accuracy, groundedness, and safety
- Own and deliver evaluation infrastructure end-to-end; mentor teams and uplevel customers on evals best practices
Must Have
- 5-7+ years building production-grade ML or GenAI systems; advanced Python
- Hands-on experience with agent evaluation: offline/on-demand regression testing and online/continuous live-traffic scoring
- LLM-as-judge proficiency: rubric design, scoring at session, response, and tool-call granularity
- Core eval metrics: task completion, tool-call accuracy, trajectory quality, groundedness, safety
- Agentic framework experience (Strands, LangGraph, CrewAI or similar)
- AWS Bedrock (Models, Guardrails) + CloudWatch; Lambda, S3, IAM fundamentals
Should Have
- Bedrock AgentCore Evaluations (on-demand SDK runner, online monitoring, A/B testing)
- Agent-as-a-Judge framework; HITL evaluation pipelines
- Benchmarking frameworks (tau-bench, SWE-bench, GAIA) and custom dataset construction
- RAG quality evaluation (faithfulness, context precision, recall)
- CI/CD integration for eval regression; observability tooling (tracing, structured logging)
- Legacy technology familiarity (Java/Spring, Sybase, DB2, or COBOL) as source context for agentic modernization or integration pattern design
- Financial services or regulated-industry experience (compliance-aware delivery: trade surveillance, data controls, audit requirements)
- Large-scale portfolio delivery experience - reusable patterns or components applied across 50+ applications, including wave planning and throughput measurement
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Evaluations Engineer in New York, NY vacancy
$197.3k - $245.6k
...year Requirements: Bachelors degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or a related field plus at least 4... ...designing distributed systems for model training, evaluation, and online inference at petabyte scale. Preferred: Experience...SuggestedFull time$200k - $250k
...Applied AI Engineer, Agents & Evaluation Compensation: $200k–$250k base + equity Location: New York City — fully onsite, 5 days per week Employment Type: Full-time Company: A fast-growing, seed-stage AI startup building claims intelligence technology...SuggestedFull timeFor contractors$197.3k - $225.1k
...we are creating responsible and reliable AI systems, changing banking for good. For... ...to build world-class applied science and engineering teams to deliver our industry leading... ...workflows, similarity search, guardrails, model evaluation, experimentation, governance, and...SuggestedFull timePart timeLocal area- Turing is seeking a Software Engineering Evaluator to create cutting-edge datasets for training and benchmarking large language models. You will... ...JavaScript (ReactJS), Java, Rust, and Go as needed. You will evaluate AI-generated code for efficiency, scalability, and reliability,...SuggestedRemote job
$80 per hour
Prolific is recruiting AI & Machine Learning Engineers to join its Expert Network, helping train and evaluate next‑gen LLMs. You’ll complete quick skill tests and, if successful, join Prolific as a participant to get paid for AI tasks requiring one hour of work or less...SuggestedRemote jobHourly payWork from homeFlexible hours- prolificacademic ltd invites experienced AI and ML engineers to join a paid expert network that supports the training and evaluation of large language models. Contributors apply deep technical judgment to assess model outputs, audit code, and provide structured human feedback...Remote jobHourly payFlexible hours
$80 per hour
Prolific is seeking AI & Machine Learning Engineers to join our Expert Network, training and evaluating the next generation of LLMs. You’ll perform paid tasks, with typical researchers paying up to $80 per hour and assignments of around one hour. Successful candidates join...Remote jobHourly payWork from homeFlexible hours- Prolific is seeking AI & Machine Learning Engineers to join its Expert Network and help train and evaluate next-generation LLMs with deep technical expertise. Successful candidates will complete a brief test and, if invited, begin as paid experts contributing to model...Remote jobFlexible hours
$229.9k - $262.4k
...AI Engineer 5 (Gen AI Platform Services: Agentic AI, Guardrails, Evaluations) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time...Full timePart timeLocal area- ...GEICO is seeking an experienced Sr. Staff AI Engineer for its AI Agent Platform to shape agent design, context strategies, and evaluation methodology at scale. You will own the intelligence layer across enterprise workflows, partnering with backend engineers to ensure...
$115k - $260k
...why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers. Sr. Staff AI Engineer, AI Agent Platform (Agent & Evaluation Harness) Why Join GEICO? GEICO is transforming how AI is built and deployed across the enterprise. As one of...Hourly payWork experience placementLocal area- OpenTrain AI, Inc. seeks a expert mathematical and statistical problem-solver to turn research workflows into self-contained terminal tasks. You will design datasets, models, constraints, and tests, and implement solutions in Python, R, Julia, C/C++, Bash, or other languages...Remote jobFor contractors
- ...are interested in this opportunity. Job Title: Akamai with AI Engineer Job Term: Contract (12 months+) Location: NYC, NY and... ...attestation mechanisms, and API-specific protections. Evaluate emerging Akamai AI security capabilities such as AI-powered detections...Contract workRelocation
- ...Role:- Applied AI Engineer – GenAI / LLMOps Location: New York, NY 10019 Duration: Contract – 12 Months Extension: Possible... ...decisions, including model selection, orchestration patterns, and evaluation strategies. Establish and enhance LLMOps practices, including...Contract work
$200k - $350k
...AI Engineer (Forward Deployed) SF/NYC / Travel / $200k-$350k + early stage equity My client is looking for a Forward Deployed... ...from prototype into reliable production use. You'll own the evaluation and reliability side too, not just the initial build, so what...$160k - $200k
...AI Engineer [Senior Associate / AVP] Founded in 1992, Cerberus is a global leader in alternative investing with approximately $71... ...or similar services, including tool use, structured outputs, evaluation and monitoring considerations. Agent frameworks and observability...$180k - $200k
...Forward Deployed AI Engineer Location: New York, NY Compensation: $180-200K + Bonus US Citizens only; No Visa Sponsorship... ...LangChain, Temporal, or LlamaIndex. Hands-on experience with AI evaluation metrics, observability tools, and responsible AI safeguards (...Local areaImmediate startVisa sponsorship- ...investment-management organization is expanding an AI technology team responsible for building the shared... ...experimentation and creating a scalable platform that enables engineers and business teams to develop, deploy, evaluate, and operate AI agents in production. This...Full time
- ...Job Title: Agentic AI Engineer Experience: 8–10 Years Employment Type: Long term Contract Location: NY/NJ/TX Job Summary... ...Azure, AWS, or Google Cloud Platform (GCP) . Implement LLM evaluation, observability, monitoring, guardrails, security, and responsible...Long term contract
$145k - $185k
...that is making a significant investment in AI as part of a major digital transformation. They are looking for a Principal AI Engineer to play a key role in defining and... ...management, caching, token optimization, evaluation, fallback strategies, latency and throughput...- ...Responsibilities Design and build agentic AI applications that can reason, plan,... ...controls, and structured outputs. Develop evaluation frameworks to measure agent accuracy,... ...environments. Partner with engineering, data, security, product, and business teams...Full time
$180k - $250k
...AI Engineer – Agentic AI / LLM Systems New York | Series A Startup $180–$250k + Equity In office (5 days) No visa sponsorship... ..., multi-agent systems, and learning workflows Build evaluation and confidence frameworks to determine when to automate vs escalate...Full timeWork at officeVisa sponsorship$200k
We're looking for an experienced AI Engineer to join Optiver's Applied AI and Platform Engineering team. In this role, you'll design,... ...including agent platforms, AI assistants, code review harnesses, evaluation frameworks, and agentic research pipelines—using Python and...Work at office$160k - $220k
...like owners, and build for long-term value. This is a hands-on AI engineering role for someone who wants to build from the ground up inside... ...use to do their work and make decisionsDesign, tune, and evaluate how agents behave in the product so the output is trustworthy...Full time$143k - $210k
CoreWeave is The Essential Cloud for AI. Built for pioneers by pioneers, CoreWeave delivers... ...Learn more at .What You'll Do:The Field Engineering organization at CoreWeave supports the... ...which the system should stop.- Build and evaluate retrieval over a large body of historical...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...That work changes the operating model an engineering organization runs on, the ways of working... ...sets the target. What we design from it is AI-native by construction. We redesign... ...libraries, standards, and governance processes.Evaluate and optimize AI performance using...Full timeWork experience placementLive inWork at officeLocal area
- ...thinking organization, apply now.We are currently seeking a Gen AI Engineer to join our team in a hybrid basis either out of our NYC or... ..., and least-privilege access.· Productionize LLMs: Build evaluation framework for open-source and foundational LLMs; implement retrieval...Work at officeRemote workFlexible hours3 days per week
$180k - $220k
...role, you will design, build, and deploy AI-powered systems that accelerate AI... ...role will partner directly with Product, Engineering, Data, Security, Legal, Compliance, Operations... ...cases end-to-end, from discovery and evaluation through launch, monitoring, and iteration...Casual workWork at officeFlexible hours- ...ability to harness advanced analytics and AI is critical to the team’s success. Our... ...strategies. We are seeking a highly skilled AI Engineer to join our Macro and Fixed Income team,... ...• Streamline Processes: Continuously evaluate the workflows of our Macro and Fixed Income...Work experience placementWork at office
$91.1k - $179.5k
Position Summary Agentic AI is moving from experimentation to production, and... ...responsibly and at scale. We're growing a team of engineers who want to work at the center of that... ..., tool integration, state management, evaluation, guardrails, and human-in-the-loop...Work at officeLocal areaVisa sponsorshipShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluations Engineer. Be the first to apply!



