The paper introduces Skill-α, a reinforcement-learning method for generating agent skills through sequential, individually evaluable edits. Its rollback reward compares downstream execution with the original and edited skills on an anchored query, providing a task-based signal for skill quality. With GPT-4o as the worker model, Skill-α improves average downstream success rates over the strongest skill-generation baseline by 3.3 points on CL-Bench and 6.7 points on tau2-bench. Ablations support the importance of both rollback rewards and progressive generation in document-to-skill and experience-to-skill settings.
No heat snapshots are available in the last 24 hours.