Practical Approaches to Evaluating AI for Public Benefit
Fall 2026Live

Practical Approaches to Evaluating AI for Public Benefit

Launch Date

September 8

15, 22, 29, Oct 6, 5 Sessions, Every Tuesday

About the Series

One AI model reportedly scored in the top 10 percent on the Uniform Bar Exam. Other models boast impressive scores on medical licensing exams, graduate-level science questions, and coding competitions. But if you're a state agency evaluating whether AI can help residents apply for benefits, assist caseworkers, review permits, answer questions accurately, or improve service delivery, those scores tell you almost nothing.

As government agencies, we need to measure how an AI tool helps us perform a task effectively, reliably, and safely over time.

Measuring success is harder than it sounds. Many organizations focus on simple metrics such as time saved, but measuring efficiency tells us too little about effectiveness: Are residents getting better outcomes? Are we working in ways that enhance democracy? Improve governance? Make the lives of those we serve and our staff better?

Complicating matters further, AI systems are not static. Performance can change both as models evolve and as staff learn new ways of using tools. A pilot that appears successful may perform very differently months later when deployed at scale.

This five-part webinar series introduces practical approaches to AI benchmarking and performance monitoring in government. 

Designed for state and local government program managers, analysts, and operational leaders, the series will provide practical tools for evaluating whether AI is delivering public value.

By the end of this series, participants will be able to:

  • Explain what AI benchmarking is, why it matters, and why it is hard
  • How agencies are identifying meaningful measures of effectiveness
  • Compare performance on real government tasks and workflows with and without AI.
  • How to use evidence to guide decisions about adoption, implementation, scaling, redesign, or retirement.
  • Distinguish between pre-adoption testing, pilot evaluation, and ongoing post-deployment monitoring.
  • Identify risks, limitations, and unacceptable errors to assess when evaluating AI systems.
Languages
Series
DateTitleDurationLed byModerated byLanguage
Oct 620262:00 PM ET
Practical Approaches to Evaluating AI for Pub…Coming Up

Principles for Public-Sector AI Evaluation

Drawing on lessons from research and practice, this session presents a practical framework for evaluating AI in government.

60 min
EN
Sept 2920262:00 PM ET
Practical Approaches to Evaluating AI for Pub…

Try Before and After You Buy

This session explores how agencies can monitor performance over time, detect emerging problems, and understand when an initially successful implementation may require adjustment or reevaluation.

60 min
EN
Sept 2220262:00 PM ET
Practical Approaches to Evaluating AI for Pub…

Comparing Humans, AI, and Human-AI Teams

Explore methods for comparing human-only, AI-only, and human-plus-AI workflows and identifying where AI adds value, where it creates risks, and where hybrid approaches perform best.

60 min
EN
Sept 1520262:00 PM ET
Practical Approaches to Evaluating AI for Pub…

Measuring What Matters

This session examines how agencies can move beyond productivity measures to define success in terms of service quality, accuracy, equity, public trust, resident and worker experience, and public outcomes.

60 min
EN
Sept 820262:00 PM ET
Practical Approaches to Evaluating AI for Pub…

Why Measuring AI Is Hard

Explore the limitations of common AI benchmarks, the differences between predictive and generative AI systems, and why performance in laboratory tests often fails to predict success in real-world settings.

60 min
EN