
Try Before and After You Buy
How can agencies assess whether an AI tool is likely to work before committing significant resources? This session covers practical methods for testing AI systems prior to procurement or deployment, including baseline comparisons, pilot design, task-based evaluation, and the identification of acceptable and unacceptable risks. Also, performance can change as models evolve, workflows shift, staff adapt their practices, and systems encounter new contexts and use cases. This session explores how agencies can monitor performance over time, detect emerging problems, and understand when an initially successful implementation may require adjustment or reevaluation.
Learning Goals