STOCKTAKE introduces a 26-week supply-chain replenishment benchmark formulated as a factored partially observable Markov decision process with six hidden factor processes. Its fair oracle uses exact per-factor Bayesian filtering and receives the same observation stream as the evaluated agent. Across 50 seeds and curated stress profiles, Claude Sonnet 5, GPT-5.4, DeepSeek-V4-Pro, and Grok 4.5 detected 84–88% of hidden failures, usually within one week, but achieved skill scores ranging from 0.62 to -0.23. The benchmark separates state estimation from control and exposes both under-response and costly over-response.
No heat snapshots are available in the last 24 hours.