The paper argues that a correct answer does not reveal why an agent succeeded after changing its information state. It introduces “success provenance” and AcquaBench, which compares CLEAN, GOLD, and SHAM substitutions across four standardized surfaces. GOLD exposes the correct target, while SHAM preserves source structure and exposure opportunity but substitutes a matched incorrect value. In D0, GOLD outperforms SHAM by 19.1–25.9 percentage points. Under distributed sufficiency in D2, the dependence persists, while coloc becomes a poor success marker, with AUROC values of 0.376 and 0.142. A supported 5.0-point CLEAN gap also shrinks to a raw GOLD difference of -0.6 points.
No heat snapshots are available in the last 24 hours.