ScrambleToolBench is an interactive terminal benchmark for testing behavioral reasoning without semantic tool schemas or documentation. Agents must discover hidden tool behavior through trial and error across a continuous curriculum, while coping with mapping drift, stochastic action failures, and temporal execution windows. Evaluations of state-of-the-art language models show that successful initial discovery does not imply robust adaptation. Under structural changes, agents often fail to trace cycles deductively and instead retain outdated beliefs or revert to exhaustive search. More test-time reasoning increases brute-force effort, while persistent memory reduces compounding errors without enabling efficient structural inference.
No heat snapshots are available in the last 24 hours.