Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents

First seen · 7/1/2026, 12:00 PMLatest activity · 7/1/2026, 12:00 PM

QVal introduces a training-free testbed for evaluating dense supervision signals in long-horizon LLM agents. Instead of judging a signal through an expensive downstream training pipeline, it measures Q-alignment: whether scores for state-action pairs rank actions consistently with the Q-values of a strong reference policy. QVal-v1.0 covers 21 methods across seven methodological families, four environments, more than 1,200 evaluation experiments, and six open-weight model backbones. According to the abstract, simple prompting baselines consistently outperform recent dense-supervision methods, with family-level performance clusters that persist across model sizes, environments, and observation modalities.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/1, 12:00 PMnot independentRepresentative
    QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents