This paper tests whether dynamic benchmarks genuinely provide contamination-free evaluation for multimodal automated fact-checking (MAFC). Studying the static AVeriTeC benchmark and a newly constructed ClaimReview2025Q4 benchmark, the authors report that 17.09%–29.30% of claims published after model knowledge cutoffs may still be contaminated. Pre-cutoff public knowledge can sometimes verify newer claims directly or through synthesis. Contamination increased Macro-F1 by up to 11.34 points and distorted system rankings, motivating stricter evaluation procedures for MAFC systems.
No heat snapshots are available in the last 24 hours.