Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models
This paper studies whether verbalized confidence can support risk-controlled human deferral for small open-weight language models. The evaluation covers 11 instruction-tuned models from three families, ranging from 0.5B to 14B parameters, on ARC-Challenge and TruthfulQA, totaling 25,168 local predictions. The authors show that strictly monotone calibration preserves the risk-coverage frontier and error-detection AUROC, while temperature scaling is fundamentally infeasible for some confidence-error patterns. A Clopper-Pearson procedure turns a 200-question calibration set into a finite-sample risk certificate under an i.i.d. deployment assumption. Platt scaling lowers ECE to 0.02, but certified autonomy remains limited.
Why it's worth reading
Small models are moving into private and offline deployments, but low calibration error does not establish safe autonomy. This paper connects confidence calibration to an auditable deferral-risk certificate.