High benchmark accuracy in molecular property prediction often conceals verbatim recall rather than scientific generalization. An audit of 22 frontier LLMs across 12 regression datasets reveals that over half the models reproduce published experimental values down to exact digits on five benchmarks. Counterintuitively, increasing reasoning depth amplifies this lookup behavior, flagging 89% more verbatim instances under identical prompts. Suppressing these memorized matches narrows the apparent performance gap among models, exposing the fragility of current AI-for-chemistry benchmarks.
No heat snapshots are available in the last 24 hours.