Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

First seen · 7/29/2026, 12:00 PMLatest activity · 7/29/2026, 12:00 PM

Agent Retrieval Bench evaluates repository context retrieval as an upstream problem for coding agents, separate from final patch correctness. It contains 427 samples across 25 repositories, including 345 positive cases, 50 natural no-gold cases, and 32 counterfactual wrong-repository controls. The corpus covers 308 base-commit snapshots, 392,000 files, and 7.9 million chunks. Qwen3-Embedding-4B leads sample-weighted MRR, Qwen3-Embedding-8B leads Recall@20, and RepoMap gives the best 8K-token budgeted context yield. Logged agent trajectories miss every gold file on 27–35% of samples.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/29, 12:00 PMnot independentRepresentative
    Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents