EviBack addresses a weakness in reinforcement learning for Agentic RAG: rollout groups in which every sample fails provide no comparative learning signal, even though some search behavior may still be useful. Its evidence-constrained Teacher backoff adds auxiliary supervision while preserving verifiable Actor rewards. The method separates evidence assessment from answer refinement, preventing reference answers from masking insufficient evidence. An automated, GPT-5.5-assisted APE pipeline produces a gated two-stage Teacher from a manually authored dual-task prompt. Across seven open-domain QA benchmarks and three Qwen3 scales, the paper reports higher F1 than Search-R1, improved single- and multi-hop macro F1, and fewer searches, duplicate queries, and forced terminations.
No heat snapshots are available in the last 24 hours.