VideoChat3 is presented as a fully open 4B-parameter video multimodal large language model designed for general, long-form, and streaming video understanding. Its efficiency relies on an Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for streaming perception. The authors also introduce three synthesized training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K. According to the abstract, VideoChat3 outperforms prior open-source models with comparable or larger parameter counts across several benchmark groups, although no concrete scores, compute figures, or ablation results are included in the supplied summary.
No heat snapshots are available in the last 24 hours.