The paper proposes Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a paradigm that transforms open-ended tasks into proxy environments with internally verifiable outcomes. Its implementation, SpyRL, uses asymmetric information and multi-agent self-play inspired by Who Is the Spy?: agents solve the same target task, then vote to identify a predetermined spy. The known spy identity provides an automatic reward signal, while successful identification is intended to correlate with response quality. The authors report improvements over existing self-improvement methods on summarization and creative writing, along with gains on mathematical reasoning. Code and models are released on GitHub.
No heat snapshots are available in the last 24 hours.