EvoAgentBench evaluates agent self-evolution through transfer of reusable procedural abilities rather than simple information retention. It covers web research, algorithmic reasoning, software engineering, and knowledge work. The benchmark extracts trace-grounded abilities from executions, canonicalizes them into operational units, and organizes related tasks with domain-specific Ability Graphs. Using a 528/267 train-test split, two scaffolds, and three backbone models, the authors report that curated ability content transfers reliably across model families, while no automatic method achieves sustained positive gains in every setting. The benchmark is intended to diagnose encoding, routing, and uptake of experience.
No heat snapshots are available in the last 24 hours.