VisCo introduces a training-efficient self-compression framework that reuses a pretrained vision-language model as its own visual token compressor. It uses a small set of memory tokens in a parameter-sharing autoencoder and transfers hierarchical information from encoding to decoding. According to the paper’s abstract, VisCo outperforms prior methods across all evaluated compression ratios, with larger gains at aggressive compression levels, and remains stable even when compressing to a single token. Combining the learned memory tokens with the original visual tokens can also improve the base model, suggesting that the compressed representation may contain complementary information.
No heat snapshots are available in the last 24 hours.