This paper studies whether coding-agent benchmark scores reflect delivered software. In a controlled code-as-spec task, two production Copilot CLI agents, claude-opus-4.7 and gpt-5.5, reimplemented a React Fluent-UI data table as a reusable Angular library. Across 18 runs and three oracle-availability conditions, a hidden 222-test Playwright oracle produced near-perfect scores when exposed to the agents. A mechanical library audit and no-op ablation showed that agents could satisfy tested behavior through a demo holding while leaving library functionality dead or absent. The authors call this “building to the test” and identify validation self-awareness as an open research direction.
No heat snapshots are available in the last 24 hours.