Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
HuggingFace Daily Papers·Varun Ursekar·Aug 5, 2026, 8:00 PM

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Papers88

HarnessOpt-Bench introduces an end-to-end benchmark for measuring whether LLMs can improve an agent harness, including prompts, tools, control flow, memory, and orchestration code. An optimizer receives a seed harness, graded feedback, and a fixed evaluation budget, then submits a candidate scored by normalized improvement on an inaccessible held-out partition. The study evaluates five frontier LLMs across four downstream tasks and 111 scored runs, finding that optimizer models differentiate more than the coding harnesses they use, while native harnesses are not consistently better.

Why it's worth reading

As agent performance increasingly depends on surrounding systems, this paper offers a common, budgeted, and auditable way to compare models that automatically improve those systems. Its findings also challenge the assumption that a model’s native coding harness is inherently superior.

Tags

AgentHarnessBenchmark自动优化评测LLMCoding Agents

Score breakdown

  • Novelty93
  • Impact88
  • Practicality82
  • Credibility86
  • Timeliness90