This paper proposes a framework for obtaining rigorous probabilistic bounds on the chance that an LLM produces harmful output for a given prompt. It applies Clopper-Pearson confidence intervals to derive probably approximately correct (PAC) bounds, and uses latent-space features to prioritize branches in the autoregressive generation tree that are more likely to be harmful. According to the abstract, the method can compute useful lower bounds even when the true harm probability is extremely small. The authors report non-trivial lower bounds on state-of-the-art LLMs, potentially enabling statistical evaluation and certification of model safety.
No heat snapshots are available in the last 24 hours.