Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports1.2

Bespoke Labs Introduces AutoResearchExam to Measure Agent Self-Improvement and Generalization

First seen · 9/10/2026, 01:08 PMLatest activity · 9/10/2026, 01:08 PM

Bespoke Labs has introduced AutoResearchExam, an evaluation benchmark designed to measure whether AI agents can conduct genuine scientific research rather than solve isolated tasks. Departing from static question answering and standard coding benchmarks, it examines an agent's capacity to adapt to unseen scientific domains, refine experimental hypotheses over successive iterations, and synthesize valid conclusions. The project establishes an empirical baseline for understanding how effectively research agents can navigate open-ended inquiry.

Event heat · last 24 hours

There are 8 persisted snapshots in the last 24 hours. Peak heat was 1.2 at 9/12, 14:00; latest heat is 1.2.

There are 8 persisted snapshots in the last 24 hours. Peak heat was 1.2 at 9/12, 14:00; latest heat is 1.2.1.20.609/12, 14:00, event heat 1.29/12, 17:00, event heat 1.29/12, 20:00, event heat 1.29/12, 23:00, event heat 1.29/13, 02:00, event heat 1.29/13, 05:00, event heat 1.29/13, 08:00, event heat 1.29/13, 11:00, event heat 1.224 hours agoNow
  1. 9/12, 14:00, event heat 1.2
  2. 9/12, 17:00, event heat 1.2
  3. 9/12, 20:00, event heat 1.2
  4. 9/12, 23:00, event heat 1.2
  5. 9/13, 02:00, event heat 1.2
  6. 9/13, 05:00, event heat 1.2
  7. 9/13, 08:00, event heat 1.2
  8. 9/13, 11:00, event heat 1.2

Reporting Timeline

  1. CommunityHacker News9/10, 01:08 PMnot independentcommunity 2 pts / 0 commentsRepresentative
    Bespoke Labs Introduces AutoResearchExam to Measure Agent Self-Improvement and Generalization