Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

First seen · 8/5/2026, 01:58 AMLatest activity · 8/5/2026, 01:58 AM

PAST-Bench evaluates whether persistent personal agents convert retained cross-session experience into better future behavior. It contains 26 scenarios and 204 episodes spanning memory, procedural reuse, information gathering, and updates, and compares matched runs with experience retention enabled or disabled. Experiments cover seven base models and four agent frameworks while checking both downstream gains and evidence for the intended save-retrieve-update pathway. The authors also introduce Hermes+, adding five targeted interventions to Hermes. According to the abstract, Hermes+ improves average gains, particularly when outdated state must be replaced, although results remain capability- and model-dependent.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/4, 04:00 AMnot independent
    PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
  2. AggregatorarXiv8/5, 01:58 AMnot independentRepresentative
    PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents