Lomekwi argues that tool-use benchmarks should separate tool discovery from tool execution. It decomposes discovery into curiosity, recognition, and efficiency: finding the components needed to construct a tool, discovering how to create it, and using the resulting tool effectively. The authors apply the framework to existing discovery tasks such as Voyager, and report that recognition performance inversely scales with model size. They also construct combinatorial games and a separate real-world-inspired environment, both of which show similar inverse-scaling behavior.
No heat snapshots are available in the last 24 hours.