AI Engineer, Evaluation
$150k - $250kDistyl AI
About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s. What We Are Looking ForAt Distyl, we build AI systems using Evaluation-Driven Development—an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter.Key ResponsibilitiesDesign and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environmentsDefine how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectivesBuild and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itselfDevelop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition aloneDefine, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolveWork closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in productionWhat We Require2+ years of software engineering experience Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application codeExperience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomesAbility to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scaleSystems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in productionAI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iterationTravel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing soWhat We OfferThe base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package100% coverage of medical, dental, and vision insurance for employee and dependentsFlexible time offRetirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resourcesComprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through CarrotComplimentary in-office lunches and snacks providedAccess to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problemsOwnership of high-impact projects across top enterprisesA mission-driven, fast-moving culture that values curiosity, pragmatism, and excellenceDistyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office..#LI-HybridWe believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.LocationSan Francisco; New YorkEmployment TypeFull timeLocation TypeHybridDepartmentEngineering
$171k - $240k
...and control spend effortlessly. Brex's AI-native automation and world-class service... ...grow your career. AI at Brex AI Engineering at Brex is redefining how businesses run... ...production environments. Experience building evaluations to measure and improve the quality,...SuggestedFull timeWork at officeRemote workWork from homeShift work$130k - $220k
...** Artificial Analysis is the leading independent AI benchmarking and insights company. They help engineers, enterprises, investors, media, and policymakers understand... ...Is** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not a...SuggestedFull timeWorldwide$150k - $210k
AI Engineer - Agentic Automation Location: Remote Compensation: $150,000 - $210,000 Join a rapidly growing company disrupting the... ...capable of autonomous decision-making and task execution - Own evaluation frameworks — accuracy, latency, cost per call, fallback...SuggestedFull timeLocal areaImmediate startRemote workFlexible hours- ...demonstrated track record of turning ambitious AI ideas into products people actually use;... ...move effortlessly between research and engineering; have shipped something extraordinary... ...of your work by regularly running evaluations and tests. Analyze production traces...SuggestedFull timeRelocation package
$180k - $400k
About the Role We're a pre-seed AI-powered HR tech startup based in San Francisco, building... .... We're looking for a mid-level AI Engineer (2-8 years of experience) who is... ...robust AI-driven user experiences. Develop evaluation and safety infrastructure to measure model...SuggestedFull timeRelocation$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering with the world’... ...of 20+ F500s. What We Are Looking ForAI Engineers build and operate production AI systems that... ...and continuously improve systems through evaluation, feedback, integration, and production...Work at office3 days per week$171k - $240k
...and control spend effortlessly. Brex’s AI-native automation and world-class service... ...you need to grow your career.AI at BrexAI Engineering at Brex is redefining how businesses run... ...gets sharper.Stand up feedback and evaluation loops that let us quickly gather product...Work at officeRemote workWork from home- We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world... ...science.Benchling is building Intelligence Engineering & Enablement, a small autonomous team... ...including MCP), memory and state management, evaluation, and observability. Make clear build vs....Work at officeLocal areaRemote workRelocationRelocation packageFlexible hours3 days per week
$120k - $200k
...you “get stuff done” end-to-end. You use AI to work smarter and solve problems... ...tooling across the spectrum: from prompt engineering and in-context learning to fine-tuned models... ...reasoning systems.Understanding of monitoring, evaluation, and iteration in production AI systems....Temporary workLocal areaWorldwide$186.5k - $328.5k
...world's leading enterprises orchestrate AI-powered work. Our vision is to expand human... ...of work with AI. About the roleAs an AI engineer at WRITER, you'll be at the forefront of... ...business needs.Contribute to the research and evaluation of emerging AI technologies, frameworks,...Full timeWork at officeLocal area$150k - $350k
...About Collate Collate is an AI document generation platform for life sciences.... ...and founder of Lever. Our AI researchers, engineers, and designers have worked at Google, Nvidia... ..., you’ll define the standards for how we evaluate, and deploy models that directly impact...Full time- ...AI has changed software development, but security hasn't caught up — until now. Corridor... ...and academia. We're hiring an AI Engineer, Product to make the AI systems that power... ...systems ~ Hands-on experience with LLM evaluation ~3+ years of experience in a software engineering...Full time
- ...About the Role Fieldguide is building AI agents for the most complex audit and advisory... ...other top-tier investors. As an AI Engineer , you'll design and build the... ...the agentic workflows, architectures, and evaluation systems that power enterprise-grade agents...Full timeWork at officeFlexible hours
$180k - $300k
...About The Role You'll own the core AI systems that power Gamma: the models, prompts... ...scale. Your job is to elevate quality, evaluate new frontier models, and push into new capabilities... ...our AI stack. You'll work closely with engineering and product to ship improvements that...Full timeWork at officeImmediate startWork from home- ...Forward Deployed AI Engineer The opportunity We are looking for a Forward Deployed AI Engineer to serve as the critical bridge between... ...domains. You understand the unique data challenges and evaluation paradigms of biological modelling. You have contributed to...Full timeShift work
- ...is a scalable, data science-first growth engine that gives B2C teams predictive clarity into... ...We're also co-building alongside leading AI companies. We're looking for an AI... ...deployment and monitoring Build and improve evaluation pipelines to measure, validate, and...Full timeShift workNight shiftWeekend work
- ...AI has changed software development. Security hasn’t caught up — until now. Corridor... ...government and academia. We’re hiring an AI Engineer to help build, experiment with, and ship... ...and maintain infrastructure for model evaluation, training, and experimentation Work closely...Full time
$150k - $250k
...Description Max AI – Stripe for Healthcare Max AI is the World’s first human-free... ...for over 10 years. And our Head of Engineering was one of the earliest engineers at Figma... ...Responsibilities Build, experiment, and evaluate AI agents and ML models in the NLP domain...Full time- ...eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and... ...’s free)! Role Overview As an AI Engineer at Fathom, you'll be hands-on with LLMs,... ...available models. Improved or created new evaluations for our existing features. By 90 Days,...Full timeWork at officeRemote work3 days per week
$7.5k
...AI Engineer Location: San Francisco, CA or Phoenix, AZ (In-Office) Partnership: EQL Tech has been exclusively retained by a high... ...regulated fintech product demands Create robust evals: build evaluation frameworks that make AI behaviour measurable, reproducible,...Full timeWork at officeRelocationVisa sponsorshipRelocation package- ...Meet Eloquent AI At Eloquent AI, we’re building the next generation of AI Operators... ...alongside world-class talent in AI, engineering, and product as we redefine the future of... ...’ performance via user simulations and evaluations. Requirements ~3+ years of experience...Full time
- ...Be one of the founding engineers at Nen, shaping the AI layer that powers automation across enterprise desktop environments at scale. The role... ...across SDK, API, and model integration layers Experience evaluating and benchmarking models with structured evals, not just...Full time
- ...revolutionizing software development with AI-powered formal verification. We've... ...About the role Join our team as an AI Engineer and help us push the boundaries of what's... ...Implement new reasoning algorithms and models Evaluate reasoning approaches, including latent...Full timeContract work
- ...Mercor's mission is to organize human intelligence to power the AI economy. We partner with leading AI labs and enterprises to... ...offices. About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of building the data...Full timeWork at officeRelocation package
$225k - $255k
...Job Description About the Role This is a founding-level AI engineering role at an early-stage B2B SaaS pricing intelligence startup... ...Your work sits at the intersection of LLM infrastructure, evaluation systems, and revenue-critical product outcomes. You'll build...Relocation- ...AI Engineer Opportunity at Goodfin Goodfin is an AI-native investment platform giving accredited investors access to pre-IPO and alternative... ...in the real world. Implement and improve RAG pipelines, evaluations, and reliability mechanisms. Monitor live AI systems,...
$180k - $250k
...AI Engineer Location: San Francisco, CA Company Stage of Funding: Seed Stage AI Startup ($6M Raised) Office Type: Onsite (5 Days Per... ...layer of the platform—building production AI agents, evaluation systems, and LLM-powered features that customers rely on every...H1bWork at officeVisa sponsorship$180k - $250k
...AI Engineer We're hiring a full-time AI Engineer to own the prompts, agents, evals, and pipelines behind user-facing features that ship... ...turn them into working prompts, agents, and pipelines. You'll evaluate them rigorously, iterate until they're production-ready, and...Full timeWork at officeRemote workRelocation$300 per month
...Forward Deployed AI Engineer As a Forward Deployed AI Engineer, you'll be embedded directly with leading scientific R&D organizations... ..., define success metrics, and support pre-sales technical evaluation. Requirements ~2+ years of professional software engineering...Full timeWork at officeRemote work- ...business problems they want to solve through AI: increasing revenue, reducing cost,... ...business value. As a Forward Deployed AI Engineer , you'll play a leading role in building... ...and tools for vector databases, evaluation, and monitoring. Our principles Show...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Engineer, Evaluation. Be the first to apply!
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- gen ai developer
- ai automation engineer




