PAST-Bench evaluates whether persistent personal agents convert retained cross-session experience into better future behavior. It contains 26 scenarios and 204 episodes spanning memory, procedural reuse, information gathering, and updates, and compares matched runs with experience retention enabled or disabled. Experiments cover seven base models and four agent frameworks while checking both downstream gains and evidence for the intended save-retrieve-update pathway. The authors also introduce Hermes+, adding five targeted interventions to Hermes. According to the abstract, Hermes+ improves average gains, particularly when outdated state must be replaced, although results remain capability- and model-dependent.
No heat snapshots are available in the last 24 hours.