Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Matthias Busch·Sep 4, 2026, 5:32 PM

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Papers82

High benchmark accuracy in molecular property prediction often conceals verbatim recall rather than scientific generalization. An audit of 22 frontier LLMs across 12 regression datasets reveals that over half the models reproduce published experimental values down to exact digits on five benchmarks. Counterintuitively, increasing reasoning depth amplifies this lookup behavior, flagging 89% more verbatim instances under identical prompts. Suppressing these memorized matches narrows the apparent performance gap among models, exposing the fragility of current AI-for-chemistry benchmarks.

Why it's worth reading

It empirically demonstrates that test-time reasoning can paradoxically amplify verbatim data retrieval, forcing a critical reassessment of benchmark reliability in AI for science.

Tags

LLMData ContaminationAI for ScienceReasoningMolecular PredictionBenchmarkingarXiv

Score breakdown

  • Novelty82
  • Impact83
  • Practicality76
  • Credibility84
  • Timeliness80