This paper argues that pairwise human feedback is an attention-limited measurement process rather than a direct readout of preference. It introduces a reduced-form model inspired by rational inattention, separating genuine reward gaps from distinctions that are difficult to detect. The authors show that passive comparisons generally cannot identify reward, attention, and default tendencies, while heterogeneous attention can cause Bradley–Terry reward models to learn misleading rankings. A Chatbot Arena case study reports a cyclic comparison component exceeding sampling noise, which no scalar reward can represent. A perceptual-comparison case study finds that response times and gaze contain gap information absent from labels.
No heat snapshots are available in the last 24 hours.