The paper extends multimodal deep research from static images to continuous video, identifying two bottlenecks: modality bias toward textual search and parametric knowledge leakage from internal memory. It introduces a decoupled perception-exploration pipeline, stage-wise tool unlocking, and a two-stage SFT plus GRPO training recipe. The authors also release Video-DR-Bench, a human-AI collaborative benchmark with 200 complex multi-hop video VQA instances. Their Video-DeepResearch-35B-A3B reports 64.0% average accuracy, above the reported scores for Claude-4.5-Sonnet, GPT-5, and Gemini 2.5 Pro.
No heat snapshots are available in the last 24 hours.