Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Engineer, Evaluation

$150k - $250k
Full-time

Distyl Ai

About Distyl AI

Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations.

We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys.

Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s.

What We Are Looking For

At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production.

 

AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems.

 

This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter.

 

Key Responsibilities

  • Design and implement evaluation frameworks that enable Evaluation-Driven Development for AI systems deployed in customer environments

  • Define how system quality is measured in each domain, ensuring that evaluation signals reflect real user needs, domain constraints, and business objectives

  • Build and maintain golden test cases and regression suites in Python, using both human-authored and AI-assisted test generation to capture critical behaviors and edge cases. These test suites are treated as first-class system components that evolve alongside the AI system itself

  • Develop and maintain evaluation pipelines—offline and online—that integrate directly into system iteration loops. Evaluation results inform prompt design, agent logic, model selection, and release readiness, ensuring that system changes are driven by measurable improvements rather than intuition alone

  • Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments. They investigate where evaluation signals diverge from real-world outcomes and refine grading approaches to maintain signal quality as systems and domains evolve

  • Work closely with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks meaningfully guide system development and deployment in production

 

What We Require

  • 2+ years of software engineering experience

  • Strong Python Engineering Skills: Write clean, maintainable Python and are comfortable building evaluation and experimentation pipelines that run in production environments. You treat evaluation code with the same rigor as application code

  • Experience with Evaluation-Driven or Experiment-Driven Development: Experience using structured evaluation or experimentation frameworks to drive system iteration, and understand the pitfalls of overfitting to metrics that don’t reflect real outcomes

  • Ability to Translate Human Judgment into Code: Work with subject matter experts to elicit high-quality judgments and encode them into test cases, scoring functions, and graders that scale

  • Systems-Oriented Mindset: Understand how evaluation interacts with prompts, agents, data, and deployment. You design evaluation systems that support fast iteration while maintaining trust and safety in production

  • AI-Native Working Style: Use AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration

  • Travel: Travel between 10-50% of the time, depending on the project, your role and level of interest in doing so

     

What We Offer

  • The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible time off

  • Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources

  • Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot

  • Complimentary in-office lunches and snacks provided

  • Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems

  • Ownership of high-impact projects across top enterprises

  • A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence

Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office. .

 

#LI-Hybrid

We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.

Vacancy posted 22 hours ago
Similar jobs that could be interesting for youBased on the AI Engineer, Evaluation in Remote vacancy
  • $161.6k - $200k

     ...To learn more, visit Hybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area) Position Summary As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and... 
    Suggested
    Local area
    Flexible hours

    Capital Rx

    New York, NY
    2 days ago
  • $40 per hour

     ...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical...  ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or... 
    Suggested
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Hartford, CT
    2 days ago
  • A cybersecurity consultancy is seeking experienced individuals to enhance AI capabilities by evaluating cybersecurity content and solving security-related challenges. You will play a pivotal role in validating AI outputs and providing critical feedback to advance cybersecurity... 
    Suggested
    Remote work
    Flexible hours

    DataAnnotation

    Phoenix, AZ
    2 days ago
  • $40 per hour

    A leading AI security firm is seeking experienced cybersecurity professionals in the United States to evaluate AI-generated security content and solve technical problems. The role offers flexible scheduling, the opportunity to work on varied projects, and hourly pay starting... 
    Suggested
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    3 days ago
  • $40 per hour

    A cybersecurity startup is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers the flexibility to work remotely while contributing to innovative security AI models. Candidates should have 2+ years... 
    Suggested
    Hourly pay
    Remote work

    DataAnnotation

    Bismarck, ND
    2 days ago
  • $40 per hour

    A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and contribute to building reliable AI tools. This remote role offers flexibility to choose projects and work hours, with pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work

    DataAnnotation

    Helena, MT
    3 days ago
  • $150k - $250k

    Slingshot Aerospace is seeking a Senior AI Engineer to join our AI and Data Science team. This role involves developing evaluation frameworks for intelligent systems in mission-critical space operations. Responsibilities include maintaining our validation SDK, designing... 
    Remote job

    Slingshot Aerospace

    New York, NY
    22 hours ago
  • A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve complex cybersecurity problems. This role is fully remote, allowing candidates to choose projects and work on their own schedule. Ideal candidates should... 
    Remote work

    DataAnnotation

    Lincoln, NE
    22 hours ago
  • $40 per hour

    A cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This role is flexible, allowing you to choose projects and work on your own schedule. Candidates should have over 2 years of hands-on cybersecurity... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Honolulu, HI
    2 days ago
  • $40 per hour

    A cybersecurity company is seeking experienced cybersecurity professionals to join their team. You will evaluate AI-generated security content, solve technical problems, and provide critical feedback to enhance AI systems. This role is remote, flexible, and offers hourly... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    3 days ago
  • $40 per hour

    A technology consulting company is seeking experienced cybersecurity professionals for a remote position. In this role, you'll evaluate AI-generated cybersecurity content, solve technical problems, and provide valuable feedback to enhance AI models. Ideal candidates should... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    22 hours ago
  • A cybersecurity technology company is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. This role offers a flexible schedule, with options for full-time or part-time remote work. Candidates should have... 
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Louisiana, MO
    22 hours ago
  • $40 per hour

    A cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. In this remote role, you will work with AI models to assess their accuracy and provide valuable feedback. Candidates should have 2+ years of... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    22 hours ago
  • A leading cybersecurity platform is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity issues. This role offers the flexibility of full-time or part-time remote work, allowing you to choose projects and set your... 
    Full time
    Part time
    Remote work

    DataAnnotation

    Topeka, KS
    2 days ago
  • $40 per hour

    A leading cybersecurity solutions provider is seeking experienced cybersecurity professionals for a remote position. You will evaluate AI-generated security content, solve technical problems, and provide essential feedback to improve AI systems. The ideal candidate will... 
    Hourly pay
    Remote work

    DataAnnotation

    Helena, MT
    3 days ago
  • $40 per hour

    A technology firm specializing in cybersecurity is seeking experienced cybersecurity professionals to help train AI models. In this role, you will evaluate AI-generated security content and solve technical problems to strengthen AI's reasoning. Candidates should have 2... 
    Hourly pay
    Remote work

    DataAnnotation

    Boston, MA
    22 hours ago
  • $40 per hour

    A cybersecurity technology company is seeking experienced professionals to evaluate AI-generated content and solve technical problems. In this remote role, candidates will work with AI systems to enhance their reasoning about real-world threats. Required qualifications... 
    Hourly pay
    Remote work

    DataAnnotation

    Lansing, MI
    22 hours ago
  • $40 per hour

    A technology company specializing in AI is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve technical problems. This remote position offers flexibility in project choice and schedule. Candidates should have over 2 years... 
    Hourly pay
    Remote work

    DataAnnotation

    Sioux Falls, SD
    2 days ago
  • A cybersecurity training company is seeking experienced cybersecurity professionals to evaluate AI-generated security content and tackle technical cybersecurity challenges. Candidates should have at least 2 years of hands-on experience in cybersecurity, along with some... 
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    New York, NY
    3 days ago
  • $40 per hour

    A leading technology firm is seeking experienced cybersecurity professionals to evaluate AI-generated security content and provide feedback for improving AI systems. This remote position offers flexibility in project selection, with an hourly pay starting at $40+. Candidates... 
    Hourly pay
    Remote work

    DataAnnotation

    Washington DC
    22 hours ago
  • $40 per hour

    A cybersecurity solutions provider is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This role allows for flexibility - working remotely from various countries including the US and requires at least 2 years of cybersecurity... 
    Hourly pay
    Remote work

    DataAnnotation

    Jackson, MS
    2 days ago
  • $40 per hour

    A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve cybersecurity problems. This role allows you to work remotely at your own pace, with projects starting at $40+ per hour. Candidates should have 2+ years... 
    Hourly pay
    Remote work

    DataAnnotation

    Wyoming, OH
    22 hours ago
  • $40 per hour

    A cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. This remote position allows you to work on your own schedule, with projects paid hourly starting at $40+. The ideal candidate... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    2 days ago
  • $40 per hour

    A cybersecurity company is seeking experienced cybersecurity professionals for a remote role. You will evaluate AI-generated security content, solve technical problems, and provide critical feedback to enhance AI systems. A minimum of 2 years of hands-on experience in cybersecurity... 
    Hourly pay
    Remote work

    DataAnnotation

    El Paso, TX
    2 days ago
  • $40 per hour

    A cybersecurity-focused technology firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This remote position offers flexible hours and the ability to choose projects, paying $40+ per hour. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Virginia, MN
    22 hours ago
  • $40 per hour

    A leader in AI training for cybersecurity is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers full-time or part-time remote work with the flexibility to choose projects and work hours. Candidates... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Madison, WI
    22 hours ago
  • $40 per hour

    A cybersecurity solutions company is seeking experienced professionals to evaluate AI-generated security content and solve technical security problems. Candidates should have over 2 years of hands-on experience in cybersecurity and coding skills, with strong writing and... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Saint Paul, MN
    2 days ago
  • $40 per hour

    A cybersecurity-focused company is seeking experienced professionals to evaluate AI-generated security content and improve AI systems. Responsibilities include assessing threats and providing technical feedback, ideal for those with 2+ years in cybersecurity roles like... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Springfield, IL
    22 hours ago
  • $40 per hour

    A cybersecurity solutions company is seeking experienced professionals to help train AI models by evaluating AI-generated security content and solving technical cybersecurity problems. This position offers flexibility to choose projects and set schedules, with compensation... 
    Hourly pay
    Remote work

    DataAnnotation

    Nevada, IA
    22 hours ago
  • $40 per hour

    A cybersecurity firm in the United States is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. You'll work directly with AI models to enhance their accuracy and improve cybersecurity tools. Ideal candidates have 2+... 
    Hourly pay
    Remote work

    DataAnnotation

    Charleston, WV
    22 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Engineer, Evaluation. Be the first to apply!