Main content start

What does it mean to measure AI?

Date
Tue July 7th 2026, 4:00pm
Location
CoDa E160
Speaker
Sanmi Koyejo, Stanford Computer Science

Benchmark rankings of AI systems are, in many respects, a solved problem, i.e., with known and controlled interventions, good benchmarks produce stable, generalizable rankings that answer which approach is better. However, trouble appears in two settings. First, with closed models, unknown interventions leave rankings stable but uninterpretable. Second, for deployment decisions, society needs scores that predict real-world outcomes, not just a rank order. Both are the same inferential failure: the score does not support the decision being made. This talk develops a measurement-science framework outlining these failures and describes a research program that treats predictive validity as an empirical criterion. It closes with open statistical problems and the institutional infrastructure needed to make evaluation a science rather than a leaderboard sport.

No conference sessions found. Try adjusting your filters and trying again.