AI Evaluation Scientist
$105k - $145kSteampunk
OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment. ContributionsImplement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics. Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs. Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations. Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements. Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports. Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety. Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments. Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability. Document evaluation processes, criteria, and results for both technical and non-technical audiences. You will contribute to the growth of our AI & Data Exploitation Practice! QualificationsAbility to hold a position of public trust with the U.S. government. Bachelor’s or Master’s degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field.2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.Proficiency in AI evaluation frameworks such as Ragas.Experience analyzing structured and unstructured data, including text, documents, and embeddings.Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.Experience working in agile or iterative development environments is a plus.Familiarity with OWASP LLM Top 10 Risks. NIH experience. Relevant certifications (helpful but not required): NIST AI RMF (AISIC)INFORMS CAPAWS/Azure/Google ML Certifications. Local to Washington, DC metro area preferred. About steampunkSteampunk relies on several factors to determine salary, including but not limited to geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this position is $105,000 to $145,000. The estimate displayed represents a typical annual salary range for this position. Annual salary is just one aspect of Steampunk’s total compensation package for employees. Learn more about additional Steampunk benefits here. Identity StatementAs part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.Steampunk is a Change Agent in the Federal contracting industry, bringing new thinking to clients in the Homeland, Federal Civilian, Health and DoD sectors. Through our Human-Centered delivery methodology, we are fundamentally changing the expectations our Federal clients have for true shared accountability in solving their toughest mission challenges. As an employee owned company, we focus on investing in our employees to enable them to do the greatest work of their careers – and rewarding them for outstanding contributions to our growth. If you want to learn more about our story, visit .Job SummaryJob ID: 7573Clearance Requirement: Public Trust
$105k - $145k
Steampunk is seeking an AI Evaluation Scientist in McLean, Virginia, to design evaluation processes for AI systems ensuring accuracy and reliability. Responsibilities include implementing evaluation frameworks and performing error analysis. Ideal candidates will have a...Suggested$77.6k - $176k
AI and ML Data ScientistThe Opportunity: As an Agentic AI Engineer and Data Scientist for military intelligence, you’re excited by the opportunity to design, develop, and deploy... ...retrieval-augmented generation pipelines, evaluation frameworks, prompt strategies, data...SuggestedFull timeContract workPart timeWork at officeLocal areaRemote work$170.87k
...working world.Consulting - Technology Consulting - AI and Quantitative Modelling - Artificial Intelligence Data Scientist (Manager) (Multiple Positions) (1731669), Ernst... ..., analyzing, and transforming data and evaluating results to make meaningful predictions and solve...SuggestedFull timeWork experience placementSummer holidayImmediate startMonday to Friday$229.9k - $262.4k
...Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking... ...with a cross-functional team of engineers, research scientists, technical program managers, and product managers to...SuggestedFull timePart timeLocal area- Nestlé IT & Digital Americas is seeking an Expert AI Data Scientist to join Nestlé USA. You will sit at the interface of business and technology... ...AI deployments. You will lead AI agent workflows, design evaluation methods, and drive adoption of advanced AI capabilities...Suggested
- Nestlé S.A.'s IT & Digital Americas team seeks an Expert AI Data Scientist to bridge business and technology, owning AI solutions end-to-end... ..., and scalable AI delivery. You will lead AI agents, design evaluation frameworks, and scale responsible AI while partnering with...
- Capital One is seeking a Senior Lead AI Engineer to design, develop, and deploy AI software components across a broad stack. The role emphasizes responsible and scalable AI systems, model evaluation, governance, and observability in production environments. You will collaborate...
$215.2k - $245.6k
Lead AI Engineer (Gen AI Platform Services: Agentic AI, Agent Guardrails, Agent Evaluation, Agent Memory) Overview At Capital One, we are creating responsible and reliable... ...cross-functional team of engineers, research scientists, technical program managers, and product...Full timePart timeLocal area$118k - $196k
...As a Managing Consultant within Guidehouse's AI and Data practice, you will help utilities and energy providers evaluate and improve energy efficiency, demand response... ...programs. Lead teams of analysts and data scientists to deliver high-quality, defensible results....Full timeTemporary workFlexible hours$161.8k - $184.6k
...Principal Data Scientist - AI Foundations, Specialist Models Data is at the center of everything we do. As a startup, we disrupted the... ...through all phases of development, from design through training, evaluation, validation, and implementation Leverage Agentic AI tools...Full timePart timeLocal areaImmediate startFlexible hours- Steampunk, Inc. is seeking a knowledgeable Data Scientist to focus on Generative AI for Federal use cases. The role involves designing advanced ML models and collaborating with engineers to implement AI solutions. Candidates should possess strong programming skills, particularly...
$204.44k - $324.99k
...APIs, and develop and test Prompt Builder templates and grounded AI experiences using Salesforce data, Data Cloud, knowledge, and... ...augmented generation (RAG) patterns Implement agent testing, evaluation, observability, and guardrails, including accuracy, hallucination...H1bLocal area$125k - $155k
...strategy analysis for our client.QualificationsElectrical Demonstrated expertise in interpreting electrical one-line diagrams and evaluating mission-critical electrical distribution systems, including utility service, switchgear, UPS, generators, transfer switches, PDUs...Work at officeLocal areaHome officeFlexible hours$82.6k - $162.8k
Position Summary We are seeking an AI Governance and Privacy Consultant who can operationalize responsible AI in real systems... ..., implement controls-as-code patterns, and stand up measurable evaluation and monitoring workflows.As a Senior Consultant, you will help...Local areaVisa sponsorship$146.83k - $192.72k
...About this team The Enterprise Data & AI team is a strategic and operational driver... ...responsibilities As an Applied AI/ML Scientist, you design, develop, test, and optimize... ...Developing experimentation frameworks and evaluation approaches to measure business impact of...Permanent employmentPart timeWork visa$274k - $376.2k
...Secure Every Identity, from AI to Human Identity is the key to unlocking the potential... ...are looking for a Principal Applied AI Scientist to take product ideas from 0 to 1 and... ...large language models Create and maintain evaluation and training datasets and workflows...Work at officeLocal areaWorldwideFlexible hours$105.4k - $207.8k
...to build the data foundations that power the next generation of AI-enabled cyber defense?If yes, then Deloitte’s Cyber team could be... ..., vector or hybrid retrieval, tool or function calling, evaluation or monitoring, prompt-injection defenses, and secure access patterns...Local areaVisa sponsorship- Jobs / AI Foundation Scientist: From Research to Real‑World AI AI Foundation Scientist: From Research to Real‑World AI Full-time About the Role A leading tech company in Virginia is seeking a passionate AI Research Scientist to help shape the future of AI-powered products...Full time
- AIToolboard is seeking an AI Foundation Scientist to advance state-of-the-art models and real-world AI applications in Virginia. You will work with large-scale datasets and modern frameworks to develop foundation models and deploy them in production. The role emphasizes...
$180k - $200k
Job DescriptionEverforth ECS is seeking a Sr. Data and AI Engineer to work Hybrid.(Typically 1-2x per week in Arlington, VA office... ...scale, complex data architectures5+ years of Data experience in evaluating, architecting, and building AI and Generative AI toolsExperience...Work at office$100k - $140k
Quantitative Researcher - Public Health, Healthcare Evaluation, & Disability About Mathematica: Mathematica applies expertise at the intersection... ...programming languages, such as R, Python, or STATA, and with AI platforms, such as ChatGPT or Microsoft Copilot Excellent...Contract workWork at officeLocal area$99k - $225k
...machine learning, and artificial intelligence (AI). If you care about moving a mission forward as... ...for the Department of Defense.As an advanced data scientist and researcher on our national security team, you’ll evaluate AI methodologies to make a real-world impact on...Full timeContract workPart timeWork at officeLocal areaRemote work- ...AIToolboard seeks a Data Scientist Principal to lead the integration of AI tooling and automated workflows for the AF/A10XC division at the Pentagon in Arlington, VA. You will drive automation initiatives, craft SOPs, and ensure DoW GenAI.mil policy compliance. The...
$146.6k - $183.25k
...AI Research Scientist - AI Biological Design The Allen Institute accelerates science for a healthier world through large-scale research designed... ...with scientists and engineers, the AI Research Scientist evaluates model performance, limitations, and scientific relevance,...Work at officeLocal areaRemote workVisa sponsorshipWork visaRelocation package- ...JOB DESCRIPTION Iron EagleX is seeking a Senior Principal AI Trainer to lead AI education, adoption, and enablement efforts across... ...with the technical AI/ML lifecycle, including inference, evaluation, and deployment concepts. ~ Experience using Python, Jupyter...Contract work
$150k - $210k
...Enterprise Knowledge (EK) is hiring for a full-time Semantic Data and AI Engineer to join our growing Knowledge and Data Services Sector.... ...pipeline architecture, embedding strategies, and response evaluation Contribute to agentic AI solution design and implementation, including...Full timeWork at officeLocal areaRemote work- /AI Research Scientist# AI Research ScientistWhatJobs DirectWashington, USFull-time## About the RoleOur client, a cutting-edge research and development... ...vision, or reinforcement learning. Design, implement, and evaluate novel AI algorithms and models. Collaborate with a team of...
- At the SEI AI Division, we conduct research in applied artificial intelligence and the engineering questions related to the practical... ...theory (supervised, unsupervised, reinforcement learning) and evaluation metrics. Hands‑on experience fine‑tuning LLMs and using frameworks...Full timeWork experience placementWork at office
$99k - $225k
AI EngineerThe Opportunity:As an AI engineer, you’ll design and deliver production-grade AI systems. You’ll own RAG pipelines, evaluation systems, and performance optimization leveraging modern AI infrastructure and tooling to build scalable, observable, and efficient services...Full timeContract workPart timeWork at officeLocal areaRemote work$110.7k - $372.9k
...struggle to navigate or afford the care they need. Deloitte has a new AI-first effort, backed by $1B in committed investment, building... ...access the right information at the right time. Reliability, evaluation & safety • Implement observability and tracing for prompts, tool...Local areaVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Scientist. Be the first to apply!


