Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Long-Horizon-Terminal-Bench: Testing Agents on Long-Horizon Terminal Tasks with Dense Reward Grading

First seen · 7/13/2026, 12:00 PMLatest activity · 7/13/2026, 12:00 PM

The paper introduces Long-Horizon-Terminal-Bench, a benchmark of 46 long-horizon terminal tasks across nine categories, including experiment reproduction, software engineering, multimodal analysis, games, and scientific computing. Tasks are decomposed into fine-grained subtasks with dense rewards and partial credit. Across 15 frontier models, a task run averages 9.9 million tokens, 231 episodes, and 85.3 minutes. The strongest tested model reaches only 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at the perfect-reward threshold of 1.0, while mean pass rates are 4.3% and 1.7%.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/13, 12:00 PMnot independentRepresentative
    Long-Horizon-Terminal-Bench: Testing Agents on Long-Horizon Terminal Tasks with Dense Reward Grading