SEED proposes a self-evolving training framework for long-horizon language-model agents. The current policy analyzes its own completed on-policy trajectories, extracts reusable natural-language skills, and then re-scores sampled actions under ordinary and skill-augmented contexts. The resulting probability shift becomes a dense token-level on-policy distillation signal, jointly optimized with outcome-based reinforcement learning. According to the abstract, experiments on text-based and vision-based agentic tasks show gains in performance, sample efficiency, and generalization to unseen scenarios. Code is available on GitHub.
No heat snapshots are available in the last 24 hours.