AI Engineer, Evaluation
$150k - $250kDistyl Ai
About Distyl AI
Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.
We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys. Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.What We Are Looking For
At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.
AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.
This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter.
Key Responsibilities
Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments
Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives
Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself
Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone
Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve
Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production
What We Require
2+ years of software engineering experience
Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code
Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomes
Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale
Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production
AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration
Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so
What We Offer
The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible time off
Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
Complimentary in-office lunches and snacks provided
Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
Ownership of high-impact projects across top enterprises
A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office. .
#LI-Hybrid
We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
$130k - $220k
...** Artificial Analysis is the leading independent AI benchmarking and insights company. They help engineers, enterprises, investors, media, and policymakers understand... ...Is** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not a...SuggestedFull timeWorldwide- ...Be one of the founding engineers at Nen, shaping the AI layer that powers automation across enterprise desktop environments at scale. The role... ...across SDK, API, and model integration layers Experience evaluating and benchmarking models with structured evals, not just...SuggestedFull time
- ...revolutionizing software development with AI-powered formal verification. We've... ...About the role Join our team as an AI Engineer and help us push the boundaries of what's... ...Implement new reasoning algorithms and models Evaluate reasoning approaches, including latent...SuggestedFull timeContract work
$180k - $300k
...About The Role You'll own the core AI systems that power Gamma: the models, prompts... ...scale. Your job is to elevate quality, evaluate new frontier models, and push into new capabilities... ...our AI stack. You'll work closely with engineering and product to ship improvements that...SuggestedFull timeWork at officeImmediate startWork from home$155k - $190k
...About Arize AI is rapidly transforming the world. As generative AI reshapes industries, teams need powerful ways... ...s where we come in. Arize AI is the leading AI & Agent Engineering observability and evaluation platform , empowering AI engineers to ship high-...SuggestedFull timeWork experience placementRemote workWork from home$300 per month
...us Edison Scientific builds and deploys AI scientist agents to accelerate science and... ...an ambitious team run by scientists and engineers from leading institutions across biology,... ...metrics, and support pre-sales technical evaluation. Requirements ~2+ years of professional...Full timeWork at officeRemote work$180k - $250k
...We're hiring a full-time AI Engineer to own the prompts, agents, evals, and pipelines behind user-facing features that ship to users.... ...turn them into working prompts, agents, and pipelines. You'll evaluate them rigorously, iterate until they're production-ready, and keep...Full timeWork at officeRemote workRelocation- ...Hiring an Applied AI Engineer to turn frontier AI research into real products. This role is for someone who understands models deeply... ...multimodal and agentic models for real-world users Train, adapt, evaluate, and improve models as needed to make the product work...Full time
$171k - $240k
...and control spend effortlessly. Brex’s AI-native automation and world-class service... ...to grow your career. AI at Brex AI Engineering at Brex is redefining how businesses run... ...gets sharper. Stand up feedback and evaluation loops that let us quickly gather product...Full timeWork at officeRemote workWork from home$150k - $350k
...About Collate Collate is an AI document generation platform for life sciences.... ...and founder of Lever. Our AI researchers, engineers, and designers have worked at Google, Nvidia... ..., you’ll define the standards for how we evaluate, and deploy models that directly impact...Full time- ...We’re hiring an AI Engineer to build the intelligence layer for the leading AI companion for language learning. You’ll own the core AI... ...if you: Think in systems. Understand how memory, latency, evaluation, and UX interact. Set a high engineering bar — nothing ships...Full time
- ...What You’ll Do Build and experiment with AI systems for code understanding, vulnerability... ...design and maintain infrastructure for model evaluation, training, and experimentation Work closely with product and engineering teams to integrate AI capabilities into Corridor...Full time
$150k - $250k
...About Distyl AI Distyl is an applied AI technology company partnering with the world... ...0+ F500s. What We Are Looking For AI Engineers build and operate production AI systems that... ...and continuously improve systems through evaluation, feedback, integration, and production...Full timeWork at officeFlexible hours3 days per week$155k - $180k
...of risk-based contracting. Arbital AI allows users to interact with complex VBC... ...actionable. We are looking for an AI Engineer to help build our next-generation conversational... ...behind them: you'll build retrieval and evaluation pipelines and the knowledge systems that...Full timeContract workWork at officeRemote workFlexible hours2 days per week- ...Forward Deployed AI Engineer The opportunity We are looking for a Forward Deployed AI Engineer to serve as the critical bridge between... ...domains. You understand the unique data challenges and evaluation paradigms of biological modelling. You have contributed to...Full timeShift work
- ...data and infrastructure layer for taste. Our goal is to end AI slop. To make AI feel right, not just be correct. We raised $18.... ...Craft agent harnesses, memory and self-improvement loops Design evaluation pipelines and synthetic data generation Create embedding and...Full time
$150k - $250k
...Description Max AI – Stripe for Healthcare Max AI is the World’s first human-free... ...for over 10 years. And our Head of Engineering was one of the earliest engineers at Figma... ...Responsibilities Build, experiment, and evaluate AI agents and ML models in the NLP domain...Full time- ...eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and... ...’s free)! Role Overview As an AI Engineer at Fathom, you'll be hands-on with LLMs,... ...available models. Improved or created new evaluations for our existing features. By 90 Days,...Full timeWork at officeRemote work3 days per week
- ...Mercor's mission is to organize human intelligence to power the AI economy. We partner with leading AI labs and enterprises to... ...offices. About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of building the data...Full timeWork at officeRelocation package
- ...Description Job Description Senior Software Engineer Job Type: Contractor (~15 hours/week)... ...experienced Senior Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex software engineering...Remote jobFor contractors
- ...Job Description Job Description San Francisco, California | Primarily On-site We are seeking an AI Evaluation Engineer – Reinforcement Learning & Agents to build the environments, evaluation systems, and supporting infrastructure used to train and assess long...
$175k - $215k
...physical dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How... ...is "real"? We are looking for a Software Engineer to build the metrics and pipelines that grade...Full timeRemote work$60 - $100 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Jack Dorsey . Position: Software Engineering, Data Science, and Systems Design Experts... ...Remote Role Responsibilities Evaluate LLM-generated responses to coding and software...Full timeContract workSummer workRemote work$204k - $259k
...of billions in simulation across 15+ U.S. states. The Planner Evaluation team works on one of the key challenges in autonomous driving:... ...the car. We are looking for experienced data-minded software engineers and data scientists to help us improve how we characterize and...Full timeRemote work$170k - $216k
...billions in simulation across 15+ U.S. states. Waymo's Release Evaluation org ensures that each version of the Waymo Driver is safe... ...objectives under resource constraints. Collaborate with other engineers, data scientists, statisticians and the leadership team to...Full timeRemote work$170k - $216k
...diverse set of sensors, enabling software engineers like you to develop multi-modal models... ...high-scale, mission-critical automation and evaluation frameworks that establish the "ultimate... ...~2+ years of experience in industrial AI applications involving the creation, maintenance...Full timeRemote work$225k - $255k
...Role We're a small, product-focused team building an AI-powered B2B pricing platform that helps companies... ...deal guidance and discount governance. As our Founding AI Engineer , you'll own the evaluation systems, feedback loops, and LLM infrastructure that allow...Relocation- ...Founding AI Engineer At Falconer, we're transforming how engineers create, access, and share knowledge with each other and their AI... ...implement backend systems in Python and/or Node.js You can evaluate tradeoffs and propose the most appropriate storage solution (SQL...Work experience placementWork at officeFlexible hours
$160k - $200k
...| Full-Time We're seeking a highly autonomous Founding AI Engineer to join a seed-stage agentic AI company applying multimodal AI... ...The core workflow is: HYPOTHESIS → EXPERIMENT → IMPLEMENT → EVALUATE → ANALYZE → ITERATE Most of the engineer's time will be...Full timeVisa sponsorship$115k - $200k
...Forward Deployed AI Engineer San Francisco, California, United States Jenn Nguyen and Friends Or refer someone About the Job Forward... ...and AI system performance measurement (precision, recall, evaluation). Understanding how LLM systems fail and how to mitigate...Work at officeVisa sponsorship2 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Engineer, Evaluation. Be the first to apply!
- machine learning ai engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai engineer contract
- machine learning ai engineer
- ai network engineer



