This paper introduces AFTER, a benchmark covering 382 realistic enterprise tasks, six professional roles, and 22 procedural skills for evaluating procedural memory in LLM agents. It tests local improvement, cross-task transfer, cross-role transfer, and cross-model generalization. According to the abstract, one refinement round improves aggregate performance by 3.7–6.7 percentage points. Skills evolved from diverse multi-model execution traces reach 73.1% cross-model test accuracy, outperforming single-model trace sources. However, some skills generalize broadly while others specialize in role-specific workflows and degrade under transfer.
No heat snapshots are available in the last 24 hours.