Mage-VL is a codec-native foundation model designed for proactive, real-time multimodal perception. Its Mage-ViT tokenizer uses codec motion vectors and residual energy to select dynamic, information-rich 16×16 patches across sparse I- and P-frames, reportedly cutting visual-token use by more than 75%. A lightweight System 1 event gate activates a causal System 2 decoder when needed. According to the abstract, Mage-VL-4B matches Qwen3-VL-4B on static tasks, improves video and spatial reasoning, reaches up to 3.5× faster wall-clock inference, and surpasses the 15B Phi-4-reasoning-vision baseline.
No heat snapshots are available in the last 24 hours.