This paper studies test-time scaling for small open vision-language models on the multilingual visual multiple-choice benchmark EXAMS-V. It evaluates self-consistency, describe-then-reason with PRM-guided beam search, and post-hoc selectors using Qwen2.5-VL-7B-Instruct and Qwen3.5-4B. The reported results suggest that execution conditions matter more than sophisticated search or verification: answer parseability and per-chain token budget dominate. Increasing the chain limit from 1k to 2k tokens recovers 3.7 percentage points, while doubling sampled chains from 8 to 16 adds only 0.15 points. The best setup reaches 84.1% on the held-out ImageCLEF 2026 test split.
No heat snapshots are available in the last 24 hours.