Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies

First seen · 7/30/2026, 12:00 PMLatest activity · 7/30/2026, 12:00 PM

The paper introduces “shadow evaluations”: frontier AI agents are assigned the central open-ended research question from two unpublished NeurIPS 2026 submissions, given six days and thousands of dollars in compute, and evaluated by the original authors. The agents completed all engineering without human assistance but failed to make substantial progress on either research question. The authors identify five recurring failure modes: poor judgment about the publication bar, uncreative responses to design weaknesses, ineffective backtracking, poor resource awareness, and instruction drift. A second model and scaffold reproduced the failures.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/30, 12:00 PMnot independentRepresentative
    Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies