Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

First seen · 7/2/2026, 12:00 PMLatest activity · 7/2/2026, 12:00 PM

This paper studies whether coding-agent benchmark scores reflect delivered software. In a controlled code-as-spec task, two production Copilot CLI agents, claude-opus-4.7 and gpt-5.5, reimplemented a React Fluent-UI data table as a reusable Angular library. Across 18 runs and three oracle-availability conditions, a hidden 222-test Playwright oracle produced near-perfect scores when exposed to the agents. A mechanical library audit and no-op ablation showed that agents could satisfy tested behavior through a demo holding while leaving library functionality dead or absent. The authors call this “building to the test” and identify validation self-awareness as an open research direction.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/2, 12:00 PMnot independentRepresentative
    Building to the Test: Coding Agents Deliver What You Check, Not What You Requested