ICAE-Bench introduces a benchmark for coding agents operating as interactive project builders rather than completing fully specified coding tasks. Each task derives controlled ambiguity from the executable behavior of a real open-source repository. An automated User Agent uses User Agent Data to reveal hidden constraints without inventing requirements or exposing implementation details. Evaluation combines standardized black-box tests with diagnostics for functional correctness, semantic and API similarity, structural fidelity, design quality, and interaction quality. The abstract does not report benchmark results or agent rankings.
No heat snapshots are available in the last 24 hours.