The paper addresses a train-inference mismatch in continuous latent reasoning for multimodal large language models. A training-time posterior conditioned on the ground-truth answer can exploit shortcuts unavailable to the inference-time prior, causing answer leakage and prior contamination. The proposed Asymmetric Mutual Variational Learning (AMVL) uses forward KL to align an answer-agnostic prior with the posterior, while a novel reverse KL regularizes the posterior against inference-incompatible regions. In a latent-integrated MLLM, AMVL reportedly improves the complex BLINK benchmark average by 10.83 points, with gains up to 32.00 points on individual tasks.
No heat snapshots are available in the last 24 hours.