Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models

First seen · 7/22/2026, 03:19 AMLatest activity · 7/22/2026, 03:19 AM

D3VL introduces a multimodal large language model framework for autonomous-driving scene understanding that jointly processes video, stereo-camera data, and LiDAR time series. The authors target a gap in current driving MLLMs, which largely focus on 2D images and video while underusing 3D sensing. D3VL uses a single, relatively simple architecture to combine 2D and 3D temporal inputs for traffic-scene and safety questions. The paper reports an 11% improvement over baseline methods on the KITTI Question-Answering dataset and introduces a Waymo QA extension covering more diverse driving conditions. Code and the dataset extension are provided on the project website.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/22, 03:19 AMnot independentRepresentative
    D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models