This paper formalizes the evidential limits of AI red-team evaluations. It defines an evaluation’s “evidential ceiling” as the maximum factor by which one result can update belief under a fixed testing budget, derives a closed-form boundary for a benchmark null result, and extends the framework to adaptive and automated red teaming through hypothesis-conditioned elicitation rates. The authors report that eight evaluation suites are adequate for high-frequency harm categories but several orders of magnitude too small for rare, catastrophic harms. The central criterion is discrimination between hypotheses, not attack success alone.
No heat snapshots are available in the last 24 hours.