This paper presents a systematic evaluation of data referencing errors (DREs) in large language models performing table tasks. Across models ranging from 1.7B to 20B parameters, the authors report that every tested model made errors such as citing incorrect values or omitting relevant entries, even when it understood the table structure. Treating data referencing as a critic improved answer accuracy by up to 12.0% through critic-based filtering and rejection sampling. A lightweight 4B-parameter critic model achieved an average F1 score of 78.2% for detecting in-distribution and out-of-distribution DREs.
No heat snapshots are available in the last 24 hours.