Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration

First seen · 7/18/2026, 11:09 PMLatest activity · 7/18/2026, 11:09 PM

This paper introduces two full-text scientific-memory benchmarks: Public AI Memory (PAIM), with 81 papers and 66 questions, and Public Transformers (PTr), with 252 papers and 98 questions. Evaluating eight memory and retrieval systems, including Theoria and a no-retrieval baseline, it finds that rankings depend heavily on ingestion granularity, raw-text preservation, retrieval budget, modality, rubric auditing, and judge choice. Graphiti leads on PAIM but retrieves 2.6 million characters per query; its advantage disappears after budget control. On PTr, sparse-dense hybrid retrieval ties for the lead among several systems.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 11:09 PMnot independentRepresentative
    Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration