Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

First seen · 7/31/2026, 12:00 PMLatest activity · 7/31/2026, 12:00 PM

This paper conducts a controlled scaling study of retrieval-augmented generation across 28 strictly nested corpus tiers spanning roughly 450x in size. It holds the questions, relevant documents, adversarial documents, reader model, and judging protocol fixed while comparing lexical, dense, graph-based, and agentic retrieval. File-System Agent performs best at the smallest shared tiers but uses 39x more query tokens at the bedrock and degrades as the search space grows. Around 10 million corpus tokens, BM25 overtakes it and leads at all larger shared tiers, approaching a 20-point margin at full scale. Dense retrieval is cheaper but less accurate, while graph RAG faces construction limits.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/31, 12:00 PMnot independentRepresentative
    BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms