AI Engineer, Evaluation
$150k - $250kDistyl Ai
About Distyl AI
Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.
We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys. Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.What We Are Looking For
At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.
AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.
This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter.
Key Responsibilities
Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives
Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself
Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone
Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve
Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production
What We Require
2+ years of software engineering experience
Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code
Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomes
Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale
Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production
AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so
What We Offer
The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible time off
Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
Complimentary in-office lunches and snacks provided
Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
Ownership of high-impact projects across top enterprises
A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office. .
#LI-Hybrid
We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
- ...Job Title AI Evaluation Engineer Location Hybrid / Remote Employment Type Full-time Job Summary We are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning...SuggestedFull timeRemote work
- ...To support innovative AI evaluation projects, the part-time AI Evaluation Engineer will create realistic developer environments, design challenging tasks for AI agents, and write tests to verify their solutions, all while working remotely. Key responsibilities Build...SuggestedPart timeRemote work
- ...greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-... ...assurance, training, and real-time call evaluation—allowing agencies to strengthen their... ...for an experienced AI Evaluation Engineer to help build and improve the next generation...SuggestedFull time
$40 per hour
A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and contribute to building reliable AI tools. This remote role offers flexibility to choose projects and work hours, with pay starting at $40+ per hour. Ideal candidates...SuggestedHourly payRemote work$40 - $180 per hour
AfterQuery is looking for experts with experience in software engineering, computer science, or app/web development to help train and evaluate AI models. This is a remote, project-based role: qualify once through a single assessment and receive ongoing project matches spanning...SuggestedRemote jobFlexible hours$50 per hour
...Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is... ...is NOT: Not data labeling. Not prompt engineering. Not writing code from scratch - the agent...Permanent employmentTemporary workPart time$70 - $80 per hour
...Role Overview Design evaluation tasks and senior-level scenarios that test AI systems used for rapid prototyping, product design, and human-AI collaborative... ...development, and cross-functional product-design-engineering collaboration. Build tasks for the Vibecoding track...Hourly payRemote work$152k - $241.5k
...believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows... ...and broad scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software...Full timeRemote work- ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model...Full time
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities include head-to-head...Remote work
$80 - $100 per hour
...verification for this role. What You'll Be Doing Design and build the coding benchmarks and evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate frontier models on real-world programming...Remote jobFull timeContract workFor contractors$118.07k - $263.16k
...the best from us. RSC2 is seeking an amazingly talented AI Engineer to join our team in Hanover, MD! Requirements In this role... ...get to provide technical advisory support for the design, evaluation, and advancement of AI-enabled applications, tools, and workflows...Full timeContract work- ...a sustainable future for local news. We are seeking an AI Engineer to design, develop, and deploy scalable LLM-powered solutions... ...Generation (RAG) systems integrating Snowflake data with LLMs. Evaluate and fine-tune foundation models via AWS Bedrock or other...Full timeTemporary workPart timeLocal area
$155k - $190k
...About Arize AI is rapidly transforming the world. As generative AI reshapes industries, teams need powerful ways... ...s where we come in. Arize AI is the leading AI & Agent Engineering observability and evaluation platform , empowering AI engineers to ship high-...Full timeWork experience placementRemote workWork from home- ...Helix AI Engineer, Robot Learning Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal... ...robot deployment . Responsibilities Design, train, evaluate, and deploy learning-based visuomotor policies for humanoid...Full time
- ...new category of enterprise software: an AI platform that changes how the world's largest... ...We don't. Tessera is a transformation engine: a governed, multi-agent platform that understands... ...search, reranking, grounding, and the evaluation that tells you whether any of it helped....Full time
- ...The AI Engineer is Lasting Change's first dedicated AI role, joining an established Data & Innovation team focused on advancing the organization... ...patterns tailored to organizational data and workflows. Evaluate, select, and integrate best-in-class LLM and AI platform...Full time
$180k - $280k
...AI Engineer Title of Role: AI Engineer Location: New York, hybrid Company Stage of Funding: Venture-Backed — Healthcare, Fintech... ...handling denials and unpaid claims. Conduct experiments to evaluate model effectiveness and iterate based on findings....Full timeWork at office- ...preparing for the future. Our services span AI Strategy, Data Intelligence, AI &... ...Intelligent Automation, Enterprise Platforms and Engineering, with a specialized focus on National... ..., retrieval-augmented generation, evaluation, and human-in-the-loop guardrails. Solid...Remote jobFull timeContract workWork at office
$200k - $300k
About Farsight Farsight is the agentic AI platform for financial services, currently... ...Ventures, supercharged by scalable engineering and AI skills from companies including Amazon... ...to polished output. Build the evaluation and quality systems behind generated deliverables...Full timeLocal areaRemote work- ...corporate office in Livonia Michigan is currently seeking an AI Engineer to join our team. The AI Engineer is responsible for building... ...data. Develop and refine prompts, system instructions, and evaluation frameworks to ensure AI outputs are accurate, consistent, and...Full timeWork at office
- ...and enterprise technology by developing AI driven solutions that improve operational... .... You will work alongside our Automation Engineers to design, develop, and implement AI models... ...for AI applications Research and evaluate emerging AI technologies, including generative...Full timeFor contractors
$110k - $140k
...developing, and deploying production-grade AI solutions including autonomous agents,... ...AI observability, guardrails, and evaluation frameworks (RAGAs, TruLens, DeepEval) to... ...code reviews and maintain high-quality engineering standards. Keep updated with advances...Full timeWork from home$155k - $200k
...the disciplined collaboration and transcendent thinking as an AI Engineer at Capstone Investment Advisors here. Responsibilities and... ...applications using Python Engage with domain experts to identify, evaluate, and execute high-value AI use cases Our future colleague...Minimum wageFull time- ...solutions. The Role We are hiring our first dedicated AI Engineer to do the same thing internally: bring our marketing, brand,... ...Set the standards for how AI is used across the company: evaluation, prompt and rule versioning, cost controls, access, and data...Permanent employmentFull timeWorldwide
- ...behavior biometrics, machine learning, and AI to stop fraud before it happens. Today,... ...fraud, account takeovers, and social engineering scams. We have raised $145M from world-class... ...engineering, fine-tuning, and rigorous evaluation frameworks to continuously optimize AI...Full timeRemote workWorldwideHome officeFlexible hours
- ...Figure is an AI robotics company developing autonomous general-purpose humanoid robots... ...autonomy. We are looking for a Helix AI Engineer, Pretraining to build large-scale... ...models into the autonomy stack Design evaluation frameworks to measure reasoning ability,...Full timeWork at office
- ...how Socure builds, deploys, and scales AI-driven identity solutions while enabling... ...Reporting to the Head of New Product Engineering, you'll join a new Internal AI Engineering... ...Operational Excellence - Leverage the evaluation harness and tracing substrate provided by...Full time
- ...based response repositories, and leverage AI to optimize workflows, processes, and... ...and structured data. Champion prompt engineering best practices across internal teams including... ...Desktop. Build and maintain internal evaluation harnesses to measure prompt quality,...Full timeContract workSecond jobWork at officeLocal area
- ...Function Chipply is hiring an Internal Forward Deployed AI Engineer to embed with our internal teams and ship AI-powered solutions... ..., output quality monitoring, and ethical use. Continuously evaluate emerging AI tools, agent frameworks, model providers, and orchestration...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Engineer, Evaluation. Be the first to apply!



