Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

First seen · 7/6/2026, 11:17 PMLatest activity · 7/6/2026, 11:17 PM

EvoAgentBench evaluates agent self-evolution through transfer of reusable procedural abilities rather than simple information retention. It covers web research, algorithmic reasoning, software engineering, and knowledge work. The benchmark extracts trace-grounded abilities from executions, canonicalizes them into operational units, and organizes related tasks with domain-specific Ability Graphs. Using a 528/267 train-test split, two scaffolds, and three backbone models, the authors report that curated ability content transfers reliably across model families, while no automatic method achieves sustained positive gains in every setting. The benchmark is intended to diagnose encoding, routing, and uptake of experience.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/6, 11:17 PMnot independentRepresentative
    EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer