This paper studies whether verbalized confidence can support risk-controlled human deferral for small open-weight language models. The evaluation covers 11 instruction-tuned models from three families, ranging from 0.5B to 14B parameters, on ARC-Challenge and TruthfulQA, totaling 25,168 local predictions. The authors show that strictly monotone calibration preserves the risk-coverage frontier and error-detection AUROC, while temperature scaling is fundamentally infeasible for some confidence-error patterns. A Clopper-Pearson procedure turns a 200-question calibration set into a finite-sample risk certificate under an i.i.d. deployment assumption. Platt scaling lowers ECE to 0.02, but certified autonomy remains limited.
No heat snapshots are available in the last 24 hours.