The paper introduces SaMer, an object-aware token merging method for multi-vector vision-language retrieval. It compresses image-side post-projector tokens into K representative centroids while preserving the original late-interaction interface. Object annotations are used only during training as a merging prior; inference requires neither ground-truth boxes nor an object detector. The vision and language backbones remain frozen, with adaptation limited to the shared projection layer. At K=64, SaMer removes more than 93% of image-side tokens, reduces ColPali storage by 16.09x, improves R@1 on Flickr30K and MSCOCO, and shows stronger phrase-level grounding than the reported compression baselines.
No heat snapshots are available in the last 24 hours.