Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Engineer, Evaluation

$150k - $250k
Full-time

Distyl Ai

About Distyl AI

Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.

We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.

Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.

What We Are Looking For

At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.

 

AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.

 

This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter.

 

Key Responsibilities

  • Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments

  • Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives

  • Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself

  • Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone

  • Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve

  • Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production

 

What We Require

  • 2+ years of software engineering experience

  • Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code

  • Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomes

  • Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale

  • Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production

  • AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration

  • Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so

     

What We Offer

  • The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible time off

  • Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources

  • Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot

  • Complimentary in-office lunches and snacks provided

  • Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems

  • Ownership of high-impact projects across top enterprises

  • A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence

Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office. .

 

#LI-Hybrid

We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Engineer, Evaluation in Remote vacancy
  •  ...Job Title AI Evaluation Engineer Location Hybrid / Remote Employment Type Full-time Job Summary We are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning... 
    Suggested
    Full time
    Remote work

    Ova Technologies

    New York, NY
    2 days ago
  •  ...To support innovative AI evaluation projects, the part-time AI Evaluation Engineer will create realistic developer environments, design challenging tasks for AI agents, and write tests to verify their solutions, all while working remotely. Key responsibilities Build... 
    Suggested
    Part time
    Remote work

    Virtual Vocations Inc

    United States
    4 hours ago
  •  ...greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-...  ...assurance, training, and real-time call evaluation—allowing agencies to strengthen their...  ...for an experienced AI Evaluation Engineer to help build and improve the next generation... 
    Suggested
    Full time

    GovWorx

    Remote
    a month ago
  • $40 per hour

    A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and contribute to building reliable AI tools. This remote role offers flexibility to choose projects and work hours, with pay starting at $40+ per hour. Ideal candidates... 
    Suggested
    Hourly pay
    Remote work

    DataAnnotation

    Helena, MT
    1 day ago
  • $40 - $180 per hour

    AfterQuery is looking for experts with experience in software engineering, computer science, or app/web development to help train and evaluate AI models. This is a remote, project-based role: qualify once through a single assessment and receive ongoing project matches spanning... 
    Suggested
    Remote job
    Flexible hours

    AfterQuery Experts

    Seattle, WA
    3 days ago
  • $50 per hour

     ...Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is...  ...is NOT: Not data labeling. Not prompt engineering. Not writing code from scratch - the agent... 
    Permanent employment
    Temporary work
    Part time

    Mindrift

    Remote
    a month ago
  • $70 - $80 per hour

     ...Role Overview Design evaluation tasks and senior-level scenarios that test AI systems used for rapid prototyping, product design, and human-AI collaborative...  ...development, and cross-functional product-design-engineering collaboration. Build tasks for the Vibecoding track... 
    Hourly pay
    Remote work

    SaidGig

    United States
    20 days ago
  • $152k - $241.5k

     ...believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows...  ...and broad scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    17 hours ago
  •  ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will build and optimize the technical foundations that power model improvement for foundation model... 
    Full time

    Innodata

    Remote
    22 days ago
  • Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities include head-to-head... 
    Remote work

    Obsidian

    New York, NY
    3 days ago
  • $80 - $100 per hour

     ...verification for this role. What You'll Be Doing Design and build the coding benchmarks and evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate frontier models on real-world programming... 
    Remote job
    Full time
    Contract work
    For contractors

    G2i

    Miami, FL
    1 day ago
  • $118.07k - $263.16k

     ...the best from us. RSC2 is seeking an amazingly talented AI Engineer to join our team in Hanover, MD! Requirements In this role...  ...get to provide technical advisory support for the design, evaluation, and advancement of AI-enabled applications, tools, and workflows... 
    Full time
    Contract work

    Rsc2, Inc.

    Remote
    1 day ago
  •  ...a sustainable future for local news.   We are seeking an AI Engineer to design, develop, and deploy scalable LLM-powered solutions...  ...Generation (RAG) systems integrating Snowflake data with LLMs. Evaluate and fine-tune foundation models via AWS Bedrock or other... 
    Full time
    Temporary work
    Part time
    Local area

    Tegna

    United States
    1 day ago
  • $155k - $190k

     ...About Arize AI is rapidly transforming the world. As generative AI reshapes industries, teams need powerful ways...  ...s where we come in. Arize AI is the leading AI & Agent Engineering observability and evaluation platform , empowering AI engineers to ship high-... 
    Full time
    Work experience placement
    Remote work
    Work from home

    Arize Ai

    Remote
    1 day ago
  •  ...Helix AI Engineer, Robot Learning Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal...  ...robot deployment . Responsibilities Design, train, evaluate, and deploy learning-based visuomotor policies for humanoid... 
    Full time

    Figure

    Remote
    1 day ago
  •  ...new category of enterprise software: an AI platform that changes how the world's largest...  ...We don't. Tessera is a transformation engine: a governed, multi-agent platform that understands...  ...search, reranking, grounding, and the evaluation that tells you whether any of it helped.... 
    Full time

    Tessera Labs

    Remote
    1 day ago
  •  ...The AI Engineer is Lasting Change's first dedicated AI role, joining an established Data & Innovation team focused on advancing the organization...  ...patterns tailored to organizational data and workflows. Evaluate, select, and integrate best-in-class LLM and AI platform... 
    Full time

    Lasting Change, Inc.

    Remote
    1 day ago
  • $180k - $280k

     ...AI Engineer Title of Role: AI Engineer Location: New York, hybrid Company Stage of Funding: Venture-Backed — Healthcare, Fintech...  ...handling denials and unpaid claims. Conduct experiments to evaluate model effectiveness and iterate based on findings.... 
    Full time
    Work at office

    Recruiting From Scratch

    Remote
    1 day ago
  •  ...preparing for the future. Our services span AI Strategy, Data Intelligence, AI &...  ...Intelligent Automation, Enterprise Platforms and Engineering, with a specialized focus on National...  ..., retrieval-augmented generation, evaluation, and human-in-the-loop guardrails. Solid... 
    Remote job
    Full time
    Contract work
    Work at office

    Anika Systems

    Remote
    1 day ago
  • $200k - $300k

    About Farsight Farsight is the agentic AI platform for financial services, currently...  ...Ventures, supercharged by scalable engineering and AI skills from companies including Amazon...  ...to polished output. Build the evaluation and quality systems behind generated deliverables... 
    Full time
    Local area
    Remote work

    Farsight Ai

    New York, NY
    1 day ago
  •  ...corporate office in Livonia Michigan is currently seeking an AI Engineer to join our team. The AI Engineer is responsible for building...  ...data. Develop and refine prompts, system instructions, and evaluation frameworks to ensure AI outputs are accurate, consistent, and... 
    Full time
    Work at office

    Sunset

    Remote
    1 day ago
  •  ...and enterprise technology by developing AI driven solutions that improve operational...  .... You will work alongside our Automation Engineers to design, develop, and implement AI models...  ...for AI applications Research and evaluate emerging AI technologies, including generative... 
    Full time
    For contractors

    Sagenet's Corporate Career Center

    Remote
    1 day ago
  • $110k - $140k

     ...developing, and deploying production-grade AI solutions including autonomous agents,...  ...AI observability, guardrails, and evaluation frameworks (RAGAs, TruLens, DeepEval) to...  ...code reviews and maintain high-quality engineering standards. Keep updated with advances... 
    Full time
    Work from home

    Confie

    Addison, TX
    1 day ago
  • $155k - $200k

     ...the disciplined collaboration and transcendent thinking as an AI Engineer at Capstone Investment Advisors here. Responsibilities and...  ...applications using Python Engage with domain experts to identify, evaluate, and execute high-value AI use cases Our future colleague... 
    Minimum wage
    Full time

    Capstone Investment Advisors

    New York, NY
    1 day ago
  •  ...solutions. The Role We are hiring our first dedicated AI Engineer to do the same thing internally: bring our marketing, brand,...  ...Set the standards for how AI is used across the company: evaluation, prompt and rule versioning, cost controls, access, and data... 
    Permanent employment
    Full time
    Worldwide

    Energyx

    Remote
    1 day ago
  •  ...behavior biometrics, machine learning, and AI to stop fraud before it happens. Today,...  ...fraud, account takeovers, and social engineering scams. We have raised $145M from world-class...  ...engineering, fine-tuning, and rigorous evaluation frameworks to continuously optimize AI... 
    Full time
    Remote work
    Worldwide
    Home office
    Flexible hours

    Sardine

    United States
    1 day ago
  •  ...Figure is an AI robotics company developing autonomous general-purpose humanoid robots...  ...autonomy. We are looking for a Helix AI Engineer, Pretraining to build large-scale...  ...models into the autonomy stack Design evaluation frameworks to measure reasoning ability,... 
    Full time
    Work at office

    Figure

    Remote
    1 day ago
  •  ...how Socure builds, deploys, and scales AI-driven identity solutions while enabling...  ...Reporting to the Head of New Product Engineering, you'll join a new Internal AI Engineering...  ...Operational Excellence - Leverage the evaluation harness and tracing substrate provided by... 
    Full time

    Socure

    Remote
    1 day ago
  •  ...based response repositories, and leverage AI to optimize workflows, processes, and...  ...and structured data. Champion prompt engineering best practices across internal teams including...  ...Desktop. Build and maintain internal evaluation harnesses to measure prompt quality,... 
    Full time
    Contract work
    Second job
    Work at office
    Local area

    Lifescience Logistics

    Remote
    1 day ago
  •  ...Function Chipply is hiring an Internal Forward Deployed AI Engineer to embed with our internal teams and ship AI-powered solutions...  ..., output quality monitoring, and ethical use. Continuously evaluate emerging AI tools, agent frameworks, model providers, and orchestration... 
    Full time

    Chipply

    United States
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Engineer, Evaluation. Be the first to apply!