From Interpretability Methods to Interpretable Models
Over a decade of explainable AI research in computer vision has assembled a mature arsenal of attribution maps, feature visualizations, and circuit analyses, yet the discipline has spent most of its energy benchmarking the tools rather than the networks themselves. This position paper proposes shifting the field's focus from interpretability methods to interpretable models. Drawing parallels to systems neuroscience, the authors advocate using existing diagnostic suites to systematically compare representations across architectures, while establishing rigorous empirical tests to measure whether independent evaluators can truly understand model behavior.
Why it's worth reading
As interpretability tools proliferate without yielding clear benchmarks for model design, this paper outlines a clear programmatic shift from evaluating methods to measuring whether models are genuinely understandable to humans.