This paper revisits how automatic harness evolution for LLM agents should be evaluated. Existing methods search over harness configurations using benchmark feedback and then report performance on the same public tasks, conflating harness improvement with additional task-level search and risking benchmark overfitting. The authors compare harness evolution with simple test-time scaling and discovery baselines under matched feedback and inference budgets. On Terminal-Bench 2.1, using GPT-5.4 and Claude Opus 4.6, harness evolution does not consistently outperform simpler search methods and shows limited transfer to held-out tasks. The work argues for fairer protocols and benchmarks for automatic harness design.
No heat snapshots are available in the last 24 hours.