AI Evaluation Scientist
$105k - $145kSteampunk
OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment. ContributionsImplement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics. Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs. Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations. Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements. Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports. Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety. Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments. Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability. Document evaluation processes, criteria, and results for both technical and non-technical audiences. You will contribute to the growth of our AI & Data Exploitation Practice! QualificationsAbility to hold a position of public trust with the U.S. government. Bachelor’s or Master’s degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field.2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.Proficiency in AI evaluation frameworks such as Ragas.Experience analyzing structured and unstructured data, including text, documents, and embeddings.Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.Experience working in agile or iterative development environments is a plus.Familiarity with OWASP LLM Top 10 Risks. NIH experience. Relevant certifications (helpful but not required): NIST AI RMF (AISIC)INFORMS CAPAWS/Azure/Google ML Certifications. Local to Washington, DC metro area preferred.About steampunkSteampunk relies on several factors to determine salary, including but not limited to geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this position is $105,000 to $145,000. The estimate displayed represents a typical annual salary range for this position. Annual salary is just one aspect of Steampunk’s total compensation package for employees. Learn more about additional Steampunk benefits here. Identity StatementAs part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.Steampunk is a Change Agent in the Federal contracting industry, bringing new thinking to clients in the Homeland, Federal Civilian, Health and DoD sectors. Through our Human-Centered delivery methodology, we are fundamentally changing the expectations our Federal clients have for true shared accountability in solving their toughest mission challenges. If you want to learn more about our story, visit .We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, or any other characteristic protected by law. Steampunk participates in the E-Verify program. Job SummaryJob ID: 7573Clearance Requirement: Public Trust
- ...Ability to Obtain Public TrustWhat You Will Do:Our consultants on the AI and Data Defense and Security team help clients maximize the... ...using Python, SQL, Spark, and cloud-native analytics platforms.Evaluate, validate, and monitor AI models for performance, explainability...SuggestedFull timeFlexible hours
$170.87k
...working world.Consulting - Technology Consulting - AI and Quantitative Modelling - Artificial Intelligence Data Scientist (Manager) (Multiple Positions) (1731669), Ernst... ..., analyzing, and transforming data and evaluating results to make meaningful predictions and solve...SuggestedFull timeWork experience placementSummer holidayImmediate startMonday to Friday$229.9k - $262.4k
## Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP)Applylocations: McLean, VA: San Francisco, CA: Cambridge, MA: San Jose, CA: New York, NYtime... ...with a cross-functional team of engineers, research scientists, technical program managers, and product managers to...SuggestedFull timePart timeLocal area$120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...hour Location: Remote Role Responsibilities Evaluate complex technical tasks using deep language expertise in Scala...SuggestedHourly payWeekly payFull timeContract workFor contractorsSummer workRemote work$36 - $72 per hour
...power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location:...SuggestedHourly payFull timeMonday to FridayFlexible hours$118k - $196k
...As a Managing Consultant within Guidehouse's AI and Data practice, you will help utilities and energy providers evaluate and improve energy efficiency, demand response... ...programs. Lead teams of analysts and data scientists to deliver high-quality, defensible results.Perform...Full timeFlexible hours$77.6k - $176k
Applied AI Health ScientistThe Opportunity:To achieve an organization’s mission, leaders... ...we need you, an experienced Applied AI Scientist for Health who can contribute expertise... ...success. What You’ll Work On: Review and evaluate AI/ML technical proposals and deliverables...Full timeContract workPart timeWork at officeLocal areaImmediate startRemote workFlexible hoursShift work$204.44k - $324.99k
...APIs, and develop and test Prompt Builder templates and grounded AI experiences using Salesforce data, Data Cloud, knowledge, and... ...augmented generation (RAG) patterns Implement agent testing, evaluation, observability, and guardrails, including accuracy, hallucination...H1bLocal area$229.9k - $262.4k
...Overview AI Engineer 5 (SDK's: Gen AI Evaluation and MCP) At Capital One, we are creating responsible and reliable AI systems, changing banking... ...with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver...Full timePart timeLocal area$150k - $210k
...Enterprise Knowledge (EK) is hiring for a full-time Semantic Data and AI Engineer to join our growing Knowledge and Data Services Sector... ...pipeline architecture, embedding strategies, and response evaluation Contribute to agentic AI solution design and implementation,...Remote jobFull timeFor contractorsH1bWork at officeLocal area$161.8k - $184.6k
...Overview Principal Data Scientist - AI Foundations, Specialist Models Data is at the center of everything we do. As a startup, we... ...through all phases of development, from design through training, evaluation, validation, and implementation Leverage Agentic AI tools...Full timePart timeLocal areaImmediate startFlexible hours$180k - $200k
Job DescriptionEverforth ECS is seeking a Sr. Data and AI Engineer to work in our Arlington, VA (Hybrid) office.(Typically 1-2x per... ...scale, complex data architectures5+ years of Data experience in evaluating, architecting, and building AI and Generative AI toolsExperience...Work at office- ...Job Description Job Description Job Title: AI Adoption Specialist Location: Langley AFB, Virginia Type: Direct Hire Work... ...direct observation of workflows and user engagement Monitor and evaluate the effectiveness of AI/RPA tools and provide recommendations...Local areaShift work
$113.4k - $245.5k
AI/ML Integration Specialist Position Description CGI is seeking an AI/ML Integration Specialist to support SAF/AQX and provide... ...Artificial Intelligence (AI) technologies. . Identify and evaluate existing AI tools for integration suitability. . Analyze...Local area- ...Position Summary * This position is contingent upon award The AI/ML Integration Specialist directs and executes all tasks... ...interoperability, API configuration, and system verification. Evaluate AI tool suitability and analyze system architectures to identify...Temporary work
- ...and prompt engineering (Preferred) 2-4+ years Experience with AI tools and software development using programming languages such... ...Data Modeling Model Tuning Prompt Engineering Model Evaluation Model Cards & Documentation Model Risk & Clearance Preferred...Work experience placement
- ...AI Consultant The AI Consultant supports the enablement, discovery, and innovation activities of a multi-party AI enablement and... ...level working sessions and workshops ~ Proven ability to develop, evaluate, and prioritize use cases in collaboration with business...Remote work
- ...AIToolboard seeks a Data Scientist Principal to lead the integration of AI tooling and automated workflows for the AF/A10XC division at the Pentagon in Arlington, VA. You will drive automation initiatives, craft SOPs, and ensure DoW GenAI.mil policy compliance. The...
$150k - $190k
...Overview Join a team where innovation meets mission. Our AI, cloud, cyber, and modernization solutions save agencies thousands... ...Build and IaC. Implement robust MLOps (experiment tracking, evaluation, bias/robustness testing, model versioning, canary/blue‑green rollouts...Full timeTemporary workImmediate startWorldwide$244.7k - $279.2k
Distinguished AI Engineer (Remote) Job Description Overview:... ...-functional team of engineers, research scientists, technical program managers, and product... ...inference, similarity search, guardrails, model evaluation, experimentation, governance, and...Full timePart timeLocal areaRemote work$146.6k - $183.25k
...AI Research Scientist - AI BioDesign The Allen Institute accelerates science for a healthier world through large-scale research designed to... ...closely with scientists and engineers, the AI Research Scientist evaluates model performance, limitations, and scientific relevance,...Work at officeLocal areaRemote workVisa sponsorshipWork visaRelocation package$159.75k - $255.6k
....Your ImpactWe are seeking a skilled and innovative Senior AI Research Scientist to join a new team focusing on agentic video and multimodal... ...problem definition and data strategy through model development, evaluation, deployment, and iteration in production environments....Work experience placementWork at officeRemote work- ...user processes. The Clinical Informaticist assesses the information and knowledge needs of healthcare professionals and patients; evaluates and refines clinical processes; and supports the development, implementation, evaluation, and continuous improvement of clinical information...Contract work
$140k - $180k
...Position Title: Senior AI Enablement Specialist Role Purpose LevelTen Energy’s mission is to accelerate the energy transition... ...— not just the hype, but what's actually relevant to LevelTen. Evaluate new tools against our security and compliance standards before...Casual workWork at officeWork from homeVisa sponsorshipFlexible hours$148.5k - $223.9k
...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive... ...experienced, driven, and creative Senior Data Scientist to join the team. We explore, invent,... ...tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout...Full time$120k - $180k
OverviewWe are seeking a Senior Data Scientist to lead the design, development, and deployment... ...engineering, model training, and model evaluation. Architect and implement advanced ML techniques... ...You will contribute to the growth of our AI & Data Exploitation Practice!...$122k - $240.5k
Position Summary Agentic AI is moving from experimentation to production, and organizations everywhere are racing to figure... ...capabilities such as orchestration, tool integration, state management, evaluation, guardrails, and human-in-the-loop controlsAt least 1 year of...Work at officeLocal areaVisa sponsorshipShift work- ...Technologies (IDT), a leading defense technology company, is seeking an AI Implementation Engineer to be part of our Warfare Systems team... ...for multi-turn user and system interactions.Implement and Evaluate AI Workflows for On-Prem systems in Air-Gapped environment: Translate...Full timeWork at officeLocal areaImmediate startShift work
$140k - $190k
OverviewWe are looking for a highly skilled AI Developer to design, build, and optimize advanced AI solutions across predictive... ..., and model improvement cycles in collaboration with Data Scientists and AI Evaluation Scientists.Design and implement advanced prompt strategies,...- ...seeking a talented and experienced Data Scientist to join our dynamic and innovative team in... ...learning algorithms.Design experiments and evaluate models to ensure accuracy, reliability,... ..., and best practices in data science and AI/ML.Do you have what it takes?Active TS/SCI...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Scientist. Be the first to apply!




