RADIO1D challenges the assumption that vision-language models require fixed, patch-based 2D features. The authors report that VLM fine-tuning makes visual representations more abstract and less spatially coherent, while image-text alignment models such as SigLIP2 develop a small set of tokens that summarize global image content. RADIO1D uses multi-teacher knowledge distillation and an autoencoder to convert images into variable-length 1D token sequences. The representation supports hierarchical summarization, scene understanding with even one token, composition-aware retrieval, and adjustable accuracy-efficiency tradeoffs in VLMs.
No heat snapshots are available in the last 24 hours.