Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Engineer

Full-time

GovWorx

AI Evaluation Engineer Location: Remote (Hybrid opportunity if live in Denver, CO) Type: Full-Time Clearance: Must have US citizenship and pass FBI fingerprint and background check in multiple states About GovWorx GovWorx is helping public safety rise to today's greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-1-1 and emergency communications centers across the country by automating quality assurance, training, and real-time call evaluation—allowing agencies to strengthen their teams and better serve their communities. or the one you already have. Position Overview We're looking for an experienced AI Evaluation Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country. This role sits at the intersection of AI engineering, prompt engineering, and data science. You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration. You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments. Key Responsibilities Design, build, and maintain automated AI evaluation pipelines for production LLM applications Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods Build offline evaluation datasets and regression testing frameworks to measure AI performance over time Analyze production AI behavior using Python, SQL, and statistical techniques to identify opportunities for improvement Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems Investigate model failures through detailed error analysis and recommend improvements to prompts, evaluation datasets, and workflows Help establish best practices for Responsible AI, evaluation methodologies, and continuous model improvement Qualifications Must-Haves Must have US citizenship and pass FBI fingerprint and background check in multiple states 3+ years of experience in software engineering, machine learning, data science, or a related technical field Experience designing evaluation metrics and interpreting AI model performance Understanding of statistical methods including hypothesis testing and experiment design Strong Python development experience Strong SQL skills with experience analyzing large datasets Experience building or supporting production LLM or Generative AI applications Experience with prompt engineering and systematic prompt evaluation Nice to Have Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker Experience building dashboards using Tableau, Sisense, Power BI, or similar tools Knowledge of Responsible AI principles and evaluation methodologies Why Join GovWorx? Help build AI systems that directly support first responders and emergency communications professionals Own AI quality, evaluation, and continuous improvement for production applications Work on cutting-edge LLM technologies and help shape the future of Responsible AI Collaborate with a high-performing team across AI, engineering, product, and data science Solve technically challenging problems with real-world impact on public safety Influence AI strategy and evaluation practices across a growing technology company

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Evaluation Engineer in United States vacancy
  •  ...Position Summary As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development... 
    Suggested
    Local area
    Flexible hours

    Capital Rx

    Denver, CO
    2 days ago
  •  ...To support the advancement of AI systems for public safety agencies, the full-time remote AI Evaluation Engineer will design and maintain automated evaluation pipelines, develop prompt engineering strategies, and analyze model performance to ensure reliability and accuracy... 
    Suggested
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    2 days ago
  • $40 per hour

    A leader in AI training for cybersecurity is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers full-time or part-time remote work with the flexibility to choose projects and work hours. Candidates... 
    Suggested
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Madison, WI
    2 days ago
  • $40 per hour

    A cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This role is flexible, allowing you to choose projects and work on your own schedule. Candidates should have over 2 years of hands-on cybersecurity... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Honolulu, HI
    3 days ago
  • $40 per hour

    A leading tech firm is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. In this remote role, you will work to enhance AI systems, requiring 2+ years of hands-on experience in cybersecurity and some coding knowledge... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Iowa, LA
    2 days ago
  • A cybersecurity consultancy is seeking experienced individuals to enhance AI capabilities by evaluating cybersecurity content and solving security-related challenges. You will play a pivotal role in validating AI outputs and providing critical feedback to advance cybersecurity... 
    Remote work
    Flexible hours

    DataAnnotation

    Phoenix, AZ
    2 days ago
  • $40 per hour

    A cybersecurity solutions company is seeking experienced professionals to evaluate AI-generated security content and solve technical security problems. Candidates should have over 2 years of hands-on experience in cybersecurity and coding skills, with strong writing and... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Saint Paul, MN
    2 days ago
  • $40 per hour

     ...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical...  ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Hartford, CT
    2 days ago
  • $40 per hour

    A cybersecurity company is seeking experienced cybersecurity professionals to join their team. You will evaluate AI-generated security content, solve technical problems, and provide critical feedback to enhance AI systems. This role is remote, flexible, and offers hourly... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    2 days ago
  • $40 per hour

    A cybersecurity startup is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers the flexibility to work remotely while contributing to innovative security AI models. Candidates should have 2+ years... 
    Hourly pay
    Remote work

    DataAnnotation

    Bismarck, ND
    2 days ago
  • $40 per hour

    A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and contribute to building reliable AI tools. This remote role offers flexibility to choose projects and work hours, with pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work

    DataAnnotation

    Helena, MT
    3 days ago
  • A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve complex cybersecurity problems. This role is fully remote, allowing candidates to choose projects and work on their own schedule. Ideal candidates should... 
    Remote work

    DataAnnotation

    Lincoln, NE
    2 days ago
  • $40 per hour

    A leading cybersecurity firm is seeking experienced professionals to join their team in evaluating AI-generated security content. You will solve technical problems and provide feedback to enhance AI capabilities related to real-world threats. The ideal candidate has over... 
    Hourly pay
    Remote work

    DataAnnotation

    Providence, RI
    2 days ago
  • $40 per hour

    A cybersecurity firm is seeking experienced cybersecurity professionals to join their team in a remote capacity. You will evaluate AI-generated security content and solve technical cybersecurity problems. The ideal candidate will have a minimum of 2 years hands-on experience... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Madison, WI
    2 days ago
  • $40 per hour

    A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This remote position offers the flexibility to choose projects and work on your own schedule, with projects starting at $40 per hour. Candidates... 
    Hourly pay
    Remote work

    DataAnnotation

    Columbia, SC
    2 days ago
  • $40 per hour

    A cybersecurity solutions provider is seeking experienced cybersecurity professionals for a REMOTE position. You will evaluate AI-generated security content, solve technical problems, and contribute to cybersecurity tools using your expertise. Candidates should have 2+... 
    Hourly pay
    Remote work

    DataAnnotation

    Oregon, WI
    2 days ago
  • $40 per hour

    A cybersecurity-focused company is looking for experienced professionals to evaluate AI-generated security content and provide feedback to improve AI systems' understanding of threats. This role, which can be full-time or part-time, allows for flexible project selection... 
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Boston, MA
    3 days ago
  • $40 per hour

    A prominent tech firm is searching for experienced cybersecurity professionals to join their remote team. In this role, you will evaluate AI-generated security content, design solutions to cybersecurity problems, and provide essential feedback for improving AI models. Candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    3 days ago
  • $40 per hour

    A leading AI training firm is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve technical problems. This role is remote, allowing you to choose your projects and work schedule. Candidates should have over 2 years of hands... 
    Hourly pay
    Remote work

    DataAnnotation

    Charleston, WV
    2 days ago
  • $40 per hour

    A leading cybersecurity solutions provider is seeking experienced cybersecurity professionals for a remote position. You will evaluate AI-generated security content, solve technical problems, and provide essential feedback to improve AI systems. The ideal candidate will... 
    Hourly pay
    Remote work

    DataAnnotation

    Helena, MT
    3 days ago
  • A leading cybersecurity platform is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity issues. This role offers the flexibility of full-time or part-time remote work, allowing you to choose projects and set your... 
    Full time
    Part time
    Remote work

    DataAnnotation

    Topeka, KS
    2 days ago
  • $130k - $220k

     ...** Artificial Analysis is the leading independent AI benchmarking and insights company. They help engineers, enterprises, investors, media, and policymakers understand...  ...Is** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not a... 
    Full time
    Worldwide

    Aurora Jobs ApS

    San Francisco, CA
    12 days ago
  • $90 per hour

    Software Engineers leverage their expertise in software development to support AI research through flexible, hourly contract work. In this role, you will evaluate AI-generated code and technical content, provide structured feedback, and contribute to enhancing AI's understanding... 
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    SaidGig

    United States
    4 days ago
  •  ...leading research accelerator for frontier AI labs and a trusted partner for global...  ...AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality...  ...Overview What Does a Typical Day Look Like? Evaluate and refine AI-generated code across... 
    For contractors
    Remote work
    Flexible hours

    Turing

    Chicago, IL
    2 days ago
  • $30 - $40 per hour

    A tech company specializing in AI seeks a Web Application Developer to evaluate and enhance AI chatbot functionality. The role involves using programming skills to solve coding challenges, assess output quality, and provide improvements. Applicants should have fluency... 
    Hourly pay
    Contract work
    Remote work
    Flexible hours

    DataAnnotation

    Springfield, IL
    3 days ago
  • $40 per hour

    A tech company specializing in AI is seeking a Web Application Developer. This remote position involves training AI models and evaluating their logic and performance. Candidates should have proficiency in at least one programming language like Python or JavaScript. Responsibilities... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    2 days ago
  • $40 per hour

    A tech firm specializing in AI training is looking for a Web Application Developer in Washington, DC. This role involves measuring the progress of AI chatbots, evaluating their outputs and logic, and providing coding challenges. Proficiency in languages such as Python... 
    Hourly pay
    Contract work
    Remote work

    DataAnnotation

    Washington DC
    2 days ago
  • $40 per hour

    A technology solutions company is seeking a Web Application Developer to improve AI models by evaluating coding outputs and performance. Candidates should be proficient in Python or JavaScript and have experience with algorithms and debugging. This remote position allows... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Providence, RI
    2 days ago
  • $150k - $250k

    Role Description As a Senior AI Engineer focused on Agentic Evaluation and Verification and Validation (V&V), you will join the AI and Data Science team within Slingshot’s Research and Development organization. You will contribute to advancing how intelligent systems are... 
    Full time
    Remote work

    Slingshot Aerospace

    Remote
    3 days ago
  • A leading AI training firm is seeking a Web Application Developer to join their team. In this role, you will train AI models, evaluate their performance, and solve coding challenges using languages like Python or JavaScript. Candidates should possess a detail-oriented... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Annapolis, MD
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!