The authors propose CritiqueDriveVLM, a three-stage framework for autonomous-driving VLMs. Multi-turn reinforcement learning with a multidimensional verifier first trains a tool-free System-2 teacher to improve logical reasoning. Latent Thought Distillation then transfers converged reasoning states into a CoT-free System-1 student. On the DriveLMM-01 benchmark, MCQ quality reportedly improves from 55.54% for the base model to 76.54%. The distilled student averages 28 generated tokens and reduces reported inference latency from 3,482 ms to 416 ms, an 88% decrease. The paper and source code are provided by the authors.
No heat snapshots are available in the last 24 hours.