This paper applies Item Response Theory (IRT) to LLM safety evaluation, fitting psychometric models to eight safety benchmarks and 192 language models. It identifies three interpretable factors—refusal strictness, truthfulness, and contextual harm—that explain most cross-model variance. Psychometrically selected items recover benchmark scores more accurately than random subsets, with roughly ten adaptive items sufficient for several benchmarks and reported evaluation-cost reductions of 97–99%. The authors also use IRT for model-level audits, including detection of naive sandbagging and changes in the model served behind APIs.
No heat snapshots are available in the last 24 hours.