This paper audits Perturbation Grounded Selection (Pgs), a training-free rule that ranks VLM answers by whether they can be reproduced under label-preserving image perturbations. Across TextVQA, MATH-Vision, MMMU, and ViLP, the supplied abstract reports apparent gains of up to 31.8 points over chain-of-thought-only majority voting. However, a format- and budget-matched control using short, no-CoT samples from the original image matches or exceeds Pgs within noise on every benchmark. The negative result suggests that decoding format, rather than perturbation consistency, explains the reported selection gains.
No heat snapshots are available in the last 24 hours.