QVal introduces a training-free testbed for evaluating dense supervision signals in long-horizon LLM agents. Instead of judging a signal through an expensive downstream training pipeline, it measures Q-alignment: whether scores for state-action pairs rank actions consistently with the Q-values of a strong reference policy. QVal-v1.0 covers 21 methods across seven methodological families, four environments, more than 1,200 evaluation experiments, and six open-weight model backbones. According to the abstract, simple prompting baselines consistently outperform recent dense-supervision methods, with family-level performance clusters that persist across model sizes, environments, and observation modalities.
No heat snapshots are available in the last 24 hours.