Practical Approaches to Evaluating AI for Public Benefit

Try Before and After You Buy

DateSeptember 29, 2026, 2:00 PM ET
Duration60 mins

How can agencies assess whether an AI tool is likely to work before committing significant resources? This session covers practical methods for testing AI systems prior to procurement or deployment, including baseline comparisons, pilot design, task-based evaluation, and the identification of acceptable and unacceptable risks. Also, performance can change as models evolve, workflows shift, staff adapt their practices, and systems encounter new contexts and use cases. This session explores how agencies can monitor performance over time, detect emerging problems, and understand when an initially successful implementation may require adjustment or reevaluation.

By the end of this workshop, participants will be able to:

  • Design practical pre-deployment evaluations that compare AI performance against current practices and establish meaningful baselines.
  • Identify acceptable and unacceptable risks through task-based testing and well-designed pilots before committing significant resources to AI adoption.
  • Develop approaches for monitoring AI performance over time and determining when changes in models, workflows, or use cases require adjustment, reevaluation, or retirement.
 
 
Patrick McLoughlin

Patrick McLoughlin

Executive Director, Maryland Benefits, State of Maryland; Former MD State Chief Data Officer

Read bio
Kathrin Frauscher

Kathrin Frauscher

Deputy Executive Director, Open Contracting Partnership

Read bio
Deborah Stine

Deborah Stine

Senior Fellow, Innovate US

Read bio

This workshop is part of the Practical Approaches to Evaluating AI for Public Benefit