Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
High benchmark accuracy in molecular property prediction often conceals verbatim recall rather than scientific generalization. An audit of 22 frontier LLMs across 12 regression datasets reveals that over half the models reproduce published experimental values down to exact digits on five benchmarks. Counterintuitively, increasing reasoning depth amplifies this lookup behavior, flagging 89% more verbatim instances under identical prompts. Suppressing these memorized matches narrows the apparent performance gap among models, exposing the fragility of current AI-for-chemistry benchmarks.
Why it's worth reading
It empirically demonstrates that test-time reasoning can paradoxically amplify verbatim data retrieval, forcing a critical reassessment of benchmark reliability in AI for science.