UltraViT targets the vision encoder, an often-overlooked bottleneck in large vision-language models deployed on edge devices. Its pyramidal architecture combines heterogeneous spatial mixers at the macro-block level while explicitly considering measured on-device latency. The authors also introduce two-stage generative pre-training: dense distillation first develops spatial features, followed by generative supervision from a capacity-mixed frozen language model. According to the supplied abstract, UltraViT outperforms encoder-centric baselines and runs nearly 1.7 times faster on-device, although hardware and benchmark details are not included in the provided metadata.
No heat snapshots are available in the last 24 hours.