This paper argues that outcome-reward reinforcement learning for search-augmented LLMs rewards correct answers but does not adequately penalize fabricated answers when retrieval fails. It proposes Abstention-Aware Reinforcement Learning (AWA-RL), which dynamically adjusts abstention rewards using query-specific prior capability estimates and ongoing on-policy observations. The authors also introduce RA-F1 to quantify the trade-off between capability and reliability. Relative to non-abstaining baselines, the paper reports up to a 10.3 percentage-point gain in precision and a 2.9-point improvement in overall RA-F1, with only a marginal reduction in raw accuracy. Code, data, and weights are released.
No heat snapshots are available in the last 24 hours.