D3VL introduces a multimodal large language model framework for autonomous-driving scene understanding that jointly processes video, stereo-camera data, and LiDAR time series. The authors target a gap in current driving MLLMs, which largely focus on 2D images and video while underusing 3D sensing. D3VL uses a single, relatively simple architecture to combine 2D and 3D temporal inputs for traffic-scene and safety questions. The paper reports an 11% improvement over baseline methods on the KITTI Question-Answering dataset and introduces a Waymo QA extension covering more diverse driving conditions. Code and the dataset extension are provided on the project website.
No heat snapshots are available in the last 24 hours.