TORINO is a plug-and-play framework for reducing visual tokens in vision-language models without fine-tuning the underlying model. It uses sparse autoencoders (SAEs) to map visual tokens into an interpretable latent space, groups tokens by shared concept activations, and applies pruning or merging within each group. The method dynamically adapts the reduction rate to image complexity instead of enforcing a fixed token budget. The abstract reports favorable efficiency-accuracy trade-offs across multiple VLM benchmarks, but provides no concrete compression ratios, latency measurements, or benchmark results.
No heat snapshots are available in the last 24 hours.