Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

First seen · 7/27/2026, 02:51 PMLatest activity · 7/27/2026, 02:51 PM

The paper argues that a correct answer does not reveal why an agent succeeded after changing its information state. It introduces “success provenance” and AcquaBench, which compares CLEAN, GOLD, and SHAM substitutions across four standardized surfaces. GOLD exposes the correct target, while SHAM preserves source structure and exposure opportunity but substitutes a matched incorrect value. In D0, GOLD outperforms SHAM by 19.1–25.9 percentage points. Under distributed sufficiency in D2, the dependence persists, while coloc becomes a poor success marker, with AUROC values of 0.376 and 0.142. A supported 5.0-point CLEAN gap also shrinks to a raw GOLD difference of -0.6 points.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/27, 02:51 PMnot independentRepresentative
    Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation