Read original
arXivJoshua Fonseca RiveraPapers89

Item Response Theory for AI Safety

This paper applies Item Response Theory (IRT) to LLM safety evaluation, fitting psychometric models to eight safety benchmarks and 192 language models. It identifies three interpretable factors—refusal strictness, truthfulness, and contextual harm—that explain most cross-model variance. Psychometrically selected items recover benchmark scores more accurately than random subsets, with roughly ten adaptive items sufficient for several benchmarks and reported evaluation-cost reductions of 97–99%. The authors also use IRT for model-level audits, including detection of naive sandbagging and changes in the model served behind APIs.

Why it's worth reading

Safety benchmarks face duplication, correlation, and evaluation-aware behavior; this work offers a concrete statistical framework for reducing test cost while auditing model behavior and API consistency.

Tags

AI safetyIRTbenchmarkingsandbaggingadaptive evaluationpsychometrics