AI Evaluation Engineer
GovWorx
AI Evaluation Engineer Location: Remote (Hybrid opportunity if live in Denver, CO) Type: Full-Time Clearance: Must have US citizenship and pass FBI fingerprint and background check in multiple states About GovWorx GovWorx is helping public safety rise to today's greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-1-1 and emergency communications centers across the country by automating quality assurance, training, and real-time call evaluation—allowing agencies to strengthen their teams and better serve their communities. or the one you already have. Position Overview We're looking for an experienced AI Evaluation Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country. This role sits at the intersection of AI engineering, prompt engineering, and data science. You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration. You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments. Key Responsibilities Design, build, and maintain automated AI evaluation pipelines for production LLM applications Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods Build offline evaluation datasets and regression testing frameworks to measure AI performance over time Analyze production AI behavior using Python, SQL, and statistical techniques to identify opportunities for improvement Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems Investigate model failures through detailed error analysis and recommend improvements to prompts, evaluation datasets, and workflows Help establish best practices for Responsible AI, evaluation methodologies, and continuous model improvement Qualifications Must-Haves Must have US citizenship and pass FBI fingerprint and background check in multiple states 3+ years of experience in software engineering, machine learning, data science, or a related technical field Experience designing evaluation metrics and interpreting AI model performance Understanding of statistical methods including hypothesis testing and experiment design Strong Python development experience Strong SQL skills with experience analyzing large datasets Experience building or supporting production LLM or Generative AI applications Experience with prompt engineering and systematic prompt evaluation Nice to Have Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker Experience building dashboards using Tableau, Sisense, Power BI, or similar tools Knowledge of Responsible AI principles and evaluation methodologies Why Join GovWorx? Help build AI systems that directly support first responders and emergency communications professionals Own AI quality, evaluation, and continuous improvement for production applications Work on cutting-edge LLM technologies and help shape the future of Responsible AI Collaborate with a high-performing team across AI, engineering, product, and data science Solve technically challenging problems with real-world impact on public safety Influence AI strategy and evaluation practices across a growing technology company
- ...Position Summary As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development...SuggestedLocal areaFlexible hours
- ...To support the advancement of AI systems for public safety agencies, the full-time remote AI Evaluation Engineer will design and maintain automated evaluation pipelines, develop prompt engineering strategies, and analyze model performance to ensure reliability and accuracy...SuggestedFull timeRemote work
$40 per hour
A leader in AI training for cybersecurity is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers full-time or part-time remote work with the flexibility to choose projects and work hours. Candidates...SuggestedHourly payFull timePart timeRemote work$40 per hour
A cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This role is flexible, allowing you to choose projects and work on your own schedule. Candidates should have over 2 years of hands-on cybersecurity...SuggestedHourly payRemote workFlexible hours$40 per hour
A leading tech firm is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. In this remote role, you will work to enhance AI systems, requiring 2+ years of hands-on experience in cybersecurity and some coding knowledge...SuggestedHourly payRemote workFlexible hours- A cybersecurity consultancy is seeking experienced individuals to enhance AI capabilities by evaluating cybersecurity content and solving security-related challenges. You will play a pivotal role in validating AI outputs and providing critical feedback to advance cybersecurity...Remote workFlexible hours
$40 per hour
A cybersecurity solutions company is seeking experienced professionals to evaluate AI-generated security content and solve technical security problems. Candidates should have over 2 years of hands-on experience in cybersecurity and coding skills, with strong writing and...Hourly payRemote workFlexible hours$40 per hour
...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical... ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or...Hourly payFull timePart timeRemote work$40 per hour
A cybersecurity company is seeking experienced cybersecurity professionals to join their team. You will evaluate AI-generated security content, solve technical problems, and provide critical feedback to enhance AI systems. This role is remote, flexible, and offers hourly...Hourly payRemote workFlexible hours$40 per hour
A cybersecurity startup is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers the flexibility to work remotely while contributing to innovative security AI models. Candidates should have 2+ years...Hourly payRemote work$40 per hour
A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and contribute to building reliable AI tools. This remote role offers flexibility to choose projects and work hours, with pay starting at $40+ per hour. Ideal candidates...Hourly payRemote work- A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve complex cybersecurity problems. This role is fully remote, allowing candidates to choose projects and work on their own schedule. Ideal candidates should...Remote work
$40 per hour
A leading cybersecurity firm is seeking experienced professionals to join their team in evaluating AI-generated security content. You will solve technical problems and provide feedback to enhance AI capabilities related to real-world threats. The ideal candidate has over...Hourly payRemote work$40 per hour
A cybersecurity firm is seeking experienced cybersecurity professionals to join their team in a remote capacity. You will evaluate AI-generated security content and solve technical cybersecurity problems. The ideal candidate will have a minimum of 2 years hands-on experience...Hourly payRemote workFlexible hours$40 per hour
A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This remote position offers the flexibility to choose projects and work on your own schedule, with projects starting at $40 per hour. Candidates...Hourly payRemote work$40 per hour
A cybersecurity solutions provider is seeking experienced cybersecurity professionals for a REMOTE position. You will evaluate AI-generated security content, solve technical problems, and contribute to cybersecurity tools using your expertise. Candidates should have 2+...Hourly payRemote work$40 per hour
A cybersecurity-focused company is looking for experienced professionals to evaluate AI-generated security content and provide feedback to improve AI systems' understanding of threats. This role, which can be full-time or part-time, allows for flexible project selection...Hourly payFull timePart timeRemote workFlexible hours$40 per hour
A prominent tech firm is searching for experienced cybersecurity professionals to join their remote team. In this role, you will evaluate AI-generated security content, design solutions to cybersecurity problems, and provide essential feedback for improving AI models. Candidates...Hourly payRemote workFlexible hours$40 per hour
A leading AI training firm is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve technical problems. This role is remote, allowing you to choose your projects and work schedule. Candidates should have over 2 years of hands...Hourly payRemote work$40 per hour
A leading cybersecurity solutions provider is seeking experienced cybersecurity professionals for a remote position. You will evaluate AI-generated security content, solve technical problems, and provide essential feedback to improve AI systems. The ideal candidate will...Hourly payRemote work- A leading cybersecurity platform is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity issues. This role offers the flexibility of full-time or part-time remote work, allowing you to choose projects and set your...Full timePart timeRemote work
$130k - $220k
...** Artificial Analysis is the leading independent AI benchmarking and insights company. They help engineers, enterprises, investors, media, and policymakers understand... ...Is** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not a...Full timeWorldwide$90 per hour
Software Engineers leverage their expertise in software development to support AI research through flexible, hourly contract work. In this role, you will evaluate AI-generated code and technical content, provide structured feedback, and contribute to enhancing AI's understanding...Remote jobHourly payContract workFlexible hours- ...leading research accelerator for frontier AI labs and a trusted partner for global... ...AI researchers who specialize in software engineering, logical reasoning, STEM, multilinguality... ...Overview What Does a Typical Day Look Like? Evaluate and refine AI-generated code across...For contractorsRemote workFlexible hours
$30 - $40 per hour
A tech company specializing in AI seeks a Web Application Developer to evaluate and enhance AI chatbot functionality. The role involves using programming skills to solve coding challenges, assess output quality, and provide improvements. Applicants should have fluency...Hourly payContract workRemote workFlexible hours$40 per hour
A tech company specializing in AI is seeking a Web Application Developer. This remote position involves training AI models and evaluating their logic and performance. Candidates should have proficiency in at least one programming language like Python or JavaScript. Responsibilities...Hourly payRemote workFlexible hours$40 per hour
A tech firm specializing in AI training is looking for a Web Application Developer in Washington, DC. This role involves measuring the progress of AI chatbots, evaluating their outputs and logic, and providing coding challenges. Proficiency in languages such as Python...Hourly payContract workRemote work$40 per hour
A technology solutions company is seeking a Web Application Developer to improve AI models by evaluating coding outputs and performance. Candidates should be proficient in Python or JavaScript and have experience with algorithms and debugging. This remote position allows...Hourly payRemote workFlexible hours$150k - $250k
Role Description As a Senior AI Engineer focused on Agentic Evaluation and Verification and Validation (V&V), you will join the AI and Data Science team within Slingshot’s Research and Development organization. You will contribute to advancing how intelligent systems are...Full timeRemote work- A leading AI training firm is seeking a Web Application Developer to join their team. In this role, you will train AI models, evaluate their performance, and solve coding challenges using languages like Python or JavaScript. Candidates should possess a detail-oriented...Hourly payRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!




