Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Attention Limited Reward Learning

First seen · 7/6/2026, 09:33 AMLatest activity · 7/6/2026, 09:33 AM

This paper argues that pairwise human feedback is an attention-limited measurement process rather than a direct readout of preference. It introduces a reduced-form model inspired by rational inattention, separating genuine reward gaps from distinctions that are difficult to detect. The authors show that passive comparisons generally cannot identify reward, attention, and default tendencies, while heterogeneous attention can cause Bradley–Terry reward models to learn misleading rankings. A Chatbot Arena case study reports a cyclic comparison component exceeding sampling noise, which no scalar reward can represent. A perceptual-comparison case study finds that response times and gaze contain gap information absent from labels.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/6, 09:33 AMnot independentRepresentative
    Attention Limited Reward Learning