Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

STOCKTAKE: Measuring the Perception–Action Gap in LLM Agents with a Fair Oracle

First seen · 7/15/2026, 05:08 PMLatest activity · 7/15/2026, 05:08 PM

STOCKTAKE introduces a 26-week supply-chain replenishment benchmark formulated as a factored partially observable Markov decision process with six hidden factor processes. Its fair oracle uses exact per-factor Bayesian filtering and receives the same observation stream as the evaluated agent. Across 50 seeds and curated stress profiles, Claude Sonnet 5, GPT-5.4, DeepSeek-V4-Pro, and Grok 4.5 detected 84–88% of hidden failures, usually within one week, but achieved skill scores ranging from 0.62 to -0.23. The benchmark separates state estimation from control and exposes both under-response and costly over-response.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/15, 05:08 PMnot independentRepresentative
    STOCKTAKE: Measuring the Perception–Action Gap in LLM Agents with a Fair Oracle