The paper introduces SESA, a Self-Evolving Skill-Augmented Agent in which a challenger poses problems, a separately parameterized solver retrieves procedural skills, and informative failures are distilled back into an evolving memory bank. Because memory changes solver trajectories, policy learning, and the challenger’s future problem distribution, task generation and skill memory co-evolve. Across seven open-domain and multi-hop QA benchmarks, SESA improves average accuracy over SSP by 1.2–3.2 points and exceeds SkillRL by 0.9 points under a unified protocol. On Qwen3, memory-free SESA-Off retains a 1.8–2.2-point gain over SSP, while inference-time retrieval adds 0.5–1.0 points.
No heat snapshots are available in the last 24 hours.