ContinualSkillBench evaluates in-context continual skill learning across five domains, each containing 100 interconnected subtasks arranged by increasing difficulty and designed for cross-task reuse. According to the abstract, sequential execution generally improves agent performance, but gains vary by model and domain. In-context learning performs comparably to explicit skill maintenance on average, implying that much of the improvement may come from adapting to prior context and feedback rather than forming reusable abstractions. Explicit skills still help selectively on tasks requiring repeatable procedures or precise outputs, while weaker models accumulate larger, more fragmented skill collections.
No heat snapshots are available in the last 24 hours.