The paper introduces Skill Self-Play (Skill-SP), a reinforcement-learning framework built around a proposer, solver, and dynamic skill controller. The proposer creates challenging tasks conditioned on sampled skills, the solver searches for solutions, and the controller updates and expands the skill library using execution feedback. The authors position skills as a middle layer between narrowly verifiable environments and open-ended self-generated tasks. They report gains on tool-use and reasoning benchmarks, including turnarounds for initially misaligned models. Code is available through the Qwen-Applications GitHub organization.
No heat snapshots are available in the last 24 hours.