HarnessOpt-Bench introduces an end-to-end benchmark for measuring whether LLMs can improve an agent harness, including prompts, tools, control flow, memory, and orchestration code. An optimizer receives a seed harness, graded feedback, and a fixed evaluation budget, then submits a candidate scored by normalized improvement on an inaccessible held-out partition. The study evaluates five frontier LLMs across four downstream tasks and 111 scored runs, finding that optimizer models differentiate more than the coding harnesses they use, while native harnesses are not consistently better.
No heat snapshots are available in the last 24 hours.