Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Engineer, Evaluation

$150k - $250k
Full-time

Distyl Ai

About Distyl AI

Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.

We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.

Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.

What We Are Looking For

At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.

 

AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.

 

This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter.

 

Key Responsibilities

  • Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments

  • Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives

  • Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself

  • Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone

  • Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve

  • Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production

 

What We Require

  • 2+ years of software engineering experience

  • Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code

  • Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomes

  • Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale

  • Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production

  • AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration

  • Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so

     

What We Offer

  • The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible time off

  • Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources

  • Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot

  • Complimentary in-office lunches and snacks provided

  • Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems

  • Ownership of high-impact projects across top enterprises

  • A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence

Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office. .

 

#LI-Hybrid

We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the AI Engineer, Evaluation in San Francisco, CA vacancy
  • $130k - $220k

     ...** Artificial Analysis is the leading independent AI benchmarking and insights company. They help engineers, enterprises, investors, media, and policymakers understand...  ...Is** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not a... 
    Suggested
    Full time
    Worldwide

    Aurora Jobs ApS

    San Francisco, CA
    more than 2 months ago
  •  ...Be one of the founding engineers at Nen, shaping the AI layer that powers automation across enterprise desktop environments at scale. The role...  ...across SDK, API, and model integration layers Experience evaluating and benchmarking models with structured evals, not just... 
    Suggested
    Full time

    Nen

    San Francisco, CA
    11 hours ago
  •  ...revolutionizing software development with AI-powered formal verification. We've...  ...About the role Join our team as an AI Engineer and help us push the boundaries of what's...  ...Implement new reasoning algorithms and models Evaluate reasoning approaches, including latent... 
    Suggested
    Full time
    Contract work

    Logical Intelligence

    San Francisco, CA
    11 hours ago
  • $180k - $300k

     ...About The Role You'll own the core AI systems that power Gamma: the models, prompts...  ...scale. Your job is to elevate quality, evaluate new frontier models, and push into new capabilities...  ...our AI stack. You'll work closely with engineering and product to ship improvements that... 
    Suggested
    Full time
    Work at office
    Immediate start
    Work from home

    Gamma

    San Francisco, CA
    11 hours ago
  • $155k - $190k

     ...About Arize AI is rapidly transforming the world. As generative AI reshapes industries, teams need powerful ways...  ...s where we come in. Arize AI is the leading AI & Agent Engineering observability and evaluation platform , empowering AI engineers to ship high-... 
    Suggested
    Full time
    Work experience placement
    Remote work
    Work from home

    Arize Ai

    San Francisco, CA
    11 hours ago
  • $300 per month

     ...us Edison Scientific builds and deploys AI scientist agents to accelerate science and...  ...an ambitious team run by scientists and engineers from leading institutions across biology,...  ...metrics, and support pre-sales technical evaluation. Requirements ~2+ years of professional... 
    Full time
    Work at office
    Remote work

    Edison Scientific Inc.

    San Francisco, CA
    11 hours ago
  • $180k - $250k

     ...We're hiring a full-time AI Engineer to own the prompts, agents, evals, and pipelines behind user-facing features that ship to users....  ...turn them into working prompts, agents, and pipelines. You'll evaluate them rigorously, iterate until they're production-ready, and keep... 
    Full time
    Work at office
    Remote work
    Relocation

    Fluency

    San Francisco, CA
    11 hours ago
  •  ...Hiring an Applied AI Engineer to turn frontier AI research into real products. This role is for someone who understands models deeply...  ...multimodal and agentic models for real-world users Train, adapt, evaluate, and improve models as needed to make the product work... 
    Full time

    Accruetalent

    San Francisco, CA
    11 hours ago
  • $171k - $240k

     ...and control spend effortlessly. Brex’s AI-native automation and world-class service...  ...to grow your career. AI at Brex AI Engineering at Brex is redefining how businesses run...  ...gets sharper. Stand up feedback and evaluation loops that let us quickly gather product... 
    Full time
    Work at office
    Remote work
    Work from home

    Brex Inc.

    San Francisco, CA
    11 hours ago
  • $150k - $350k

     ...About Collate   Collate is an AI document generation platform for life sciences....  ...and founder of Lever. Our AI researchers, engineers, and designers have worked at Google, Nvidia...  ..., you’ll define the standards for how we evaluate, and deploy models that directly impact... 
    Full time

    Collate

    San Francisco, CA
    11 hours ago
  •  ...We’re hiring an AI Engineer to build the intelligence layer for the leading AI companion for language learning. You’ll own the core AI...  ...if you: Think in systems. Understand how memory, latency, evaluation, and UX interact. Set a high engineering bar — nothing ships... 
    Full time

    Pingo پینگو

    San Francisco, CA
    11 hours ago
  •  ...What You’ll Do Build and experiment with AI systems for code understanding, vulnerability...  ...design and maintain infrastructure for model evaluation, training, and experimentation Work closely with product and engineering teams to integrate AI capabilities into Corridor... 
    Full time

    Corridor

    San Francisco, CA
    11 hours ago
  • $150k - $250k

     ...About Distyl AI Distyl is an applied AI technology company partnering with the world...  ...0+ F500s. What We Are Looking For AI Engineers build and operate production AI systems that...  ...and continuously improve systems through evaluation, feedback, integration, and production... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Distyl Ai

    San Francisco, CA
    11 hours ago
  • $155k - $180k

     ...of risk-based contracting. Arbital AI allows users to interact with complex VBC...  ...actionable.  We are looking for an AI Engineer to help build our next-generation conversational...  ...behind them: you'll build retrieval and evaluation pipelines and the knowledge systems that... 
    Full time
    Contract work
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Arbital Health

    San Francisco, CA
    11 hours ago
  •  ...Forward Deployed AI Engineer The opportunity We are looking for a Forward Deployed AI Engineer to serve as the critical bridge between...  ...domains. You understand the unique data challenges and evaluation paradigms of biological modelling. You have contributed to... 
    Full time
    Shift work

    Latent Labs

    San Francisco, CA
    11 hours ago
  •  ...data and infrastructure layer for taste. Our goal is to end AI slop. To make AI feel right, not just be correct. We raised $18....  ...Craft agent harnesses, memory and self-improvement loops Design evaluation pipelines and synthetic data generation Create embedding and... 
    Full time

    Taste Labs

    San Francisco, CA
    11 hours ago
  • $150k - $250k

     ...Description Max AI – Stripe for Healthcare Max AI is the World’s first human-free...  ...for over 10 years. And our Head of Engineering was one of the earliest engineers at Figma...  ...Responsibilities Build, experiment, and evaluate AI agents and ML models in the NLP domain... 
    Full time

    Maxai

    San Francisco, CA
    11 hours ago
  •  ...eliminate the needless overhead of meetings. Our AI assistant captures, summarizes, and...  ...’s free)! Role Overview As an AI Engineer at Fathom, you'll be hands-on with LLMs,...  ...available models. Improved or created new evaluations for our existing features. By 90 Days,... 
    Full time
    Work at office
    Remote work
    3 days per week

    Fathom

    San Francisco, CA
    11 hours ago
  •  ...Mercor's mission is to organize human intelligence to power the AI economy. We partner with leading AI labs and enterprises to...  ...offices. About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of building the data... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    11 hours ago
  •  ...Description Job Description Senior Software Engineer Job Type: Contractor (~15 hours/week)...  ...experienced Senior Software Engineers to support an AI training project by creating reinforcement learning environments that evaluate AI models on complex software engineering... 
    Remote job
    For contractors

    YO AI Labs

    San Francisco, CA
    26 days ago
  •  ...Job Description Job Description San Francisco, California | Primarily On-site We are seeking an AI Evaluation Engineer – Reinforcement Learning & Agents to build the environments, evaluation systems, and supporting infrastructure used to train and assess long... 

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    12 days ago
  • $175k - $215k

     ...physical dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How...  ...is "real"? We are looking for a Software Engineer to build the metrics and pipelines that grade... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    11 hours ago
  • $60 - $100 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Jack Dorsey . Position: Software Engineering, Data Science, and Systems Design Experts...  ...Remote Role Responsibilities Evaluate LLM-generated responses to coding and software... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    11 hours ago
  • $204k - $259k

     ...of billions in simulation across 15+ U.S. states. The Planner Evaluation team works on one of the key challenges in autonomous driving:...  ...the car. We are looking for experienced data-minded software engineers and data scientists to help us improve how we characterize and... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    11 hours ago
  • $170k - $216k

     ...billions in simulation across 15+ U.S. states. Waymo's Release Evaluation org ensures that each version of the Waymo Driver is safe...  ...objectives under resource constraints. Collaborate with other engineers, data scientists, statisticians and the leadership team to... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    11 hours ago
  • $170k - $216k

     ...diverse set of sensors, enabling software engineers like you to develop multi-modal models...  ...high-scale, mission-critical automation and evaluation frameworks that establish the "ultimate...  ...~2+ years of experience in industrial AI applications involving the creation, maintenance... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    11 hours ago
  • $225k - $255k

     ...Role We're a small, product-focused team building an AI-powered B2B pricing platform that helps companies...  ...deal guidance and discount governance. As our Founding AI Engineer , you'll own the evaluation systems, feedback loops, and LLM infrastructure that allow... 
    Relocation

    Clera

    San Francisco, CA
    11 hours ago
  •  ...Founding AI Engineer At Falconer, we're transforming how engineers create, access, and share knowledge with each other and their AI...  ...implement backend systems in Python and/or Node.js You can evaluate tradeoffs and propose the most appropriate storage solution (SQL... 
    Work experience placement
    Work at office
    Flexible hours

    Falconer AI

    San Francisco, CA
    4 days ago
  • $160k - $200k

     ...| Full-Time We're seeking a highly autonomous Founding AI Engineer to join a seed-stage agentic AI company applying multimodal AI...  ...The core workflow is: HYPOTHESIS → EXPERIMENT → IMPLEMENT → EVALUATE → ANALYZE → ITERATE Most of the engineer's time will be... 
    Full time
    Visa sponsorship

    JeffreyM Consulting

    San Francisco, CA
    1 day ago
  • $115k - $200k

     ...Forward Deployed AI Engineer San Francisco, California, United States Jenn Nguyen and Friends Or refer someone About the Job Forward...  ...and AI system performance measurement (precision, recall, evaluation). Understanding how LLM systems fail and how to mitigate... 
    Work at office
    Visa sponsorship
    2 days per week

    Jenn Nguyen and Friends

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Engineer, Evaluation. Be the first to apply!