Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Item Response Theory for AI Safety

First seen · 8/6/2026, 01:25 AMLatest activity · 8/6/2026, 01:25 AM

This paper applies Item Response Theory (IRT) to LLM safety evaluation, fitting psychometric models to eight safety benchmarks and 192 language models. It identifies three interpretable factors—refusal strictness, truthfulness, and contextual harm—that explain most cross-model variance. Psychometrically selected items recover benchmark scores more accurately than random subsets, with roughly ten adaptive items sufficient for several benchmarks and reported evaluation-cost reductions of 97–99%. The authors also use IRT for model-level audits, including detection of naive sandbagging and changes in the model served behind APIs.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv8/6, 01:25 AMnot independentRepresentative
    Item Response Theory for AI Safety