Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingNewsWatching0 independent reports0

VideoChat3: A Fully Open, Efficient, and Generalist Video MLLM

First seen · 7/17/2026, 12:00 PMLatest activity · 7/17/2026, 12:00 PM

VideoChat3 is presented as a fully open 4B-parameter video multimodal large language model designed for general, long-form, and streaming video understanding. Its efficiency relies on an Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for streaming perception. The authors also introduce three synthesized training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K. According to the abstract, VideoChat3 outperforms prior open-source models with comparable or larger parameter counts across several benchmark groups, although no concrete scores, compute figures, or ablation results are included in the supplied summary.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/17, 12:00 PMnot independentRepresentative
    VideoChat3: A Fully Open, Efficient, and Generalist Video MLLM