OpenCoF studies Chain-of-Frame (CoF) reasoning, where intermediate reasoning unfolds across temporally connected video frames. The framework includes OpenCoF-17K, a dataset covering 11 reasoning task families, and Wan-CoF, a model fine-tuned from Wan2.2-I2V-A14B. The paper reports substantial gains over the baseline on four video reasoning benchmarks. It also introduces visual and textual reasoning tokens to represent low-level visual cues and high-level semantic priors, then analyzes their behavior across model depth, denoising steps, spatial structure, and temporal structure. The dataset, model, and code are released for further research.
No heat snapshots are available in the last 24 hours.