This paper argues that exposing logged ground-truth future trajectories to teacher models during chain-of-thought annotation creates trajectory anchoring bias: models rationalize observed outcomes instead of inferring decisions from scene evidence, leading to less causally faithful reasoning and more severe hallucinations in difficult scenes. It introduces Autonomous-Driving Multiple-Choice Questions (AD-MCQ), which frames planning as selection among explicit candidate trajectories, and Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR), which withholds future trajectories until after the decision. The abstract reports improved autonomous-driving reasoning while preserving or improving general visual capabilities.
No heat snapshots are available in the last 24 hours.