Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

First seen · 7/7/2026, 12:00 PMLatest activity · 7/7/2026, 12:00 PM

The paper introduces SaMer, an object-aware token merging method for multi-vector vision-language retrieval. It compresses image-side post-projector tokens into K representative centroids while preserving the original late-interaction interface. Object annotations are used only during training as a merging prior; inference requires neither ground-truth boxes nor an object detector. The vision and language backbones remain frozen, with adaptation limited to the shared projection layer. At K=64, SaMer removes more than 93% of image-side tokens, reduces ColPali storage by 16.09x, improves R@1 on Flickr30K and MSCOCO, and shows stronger phrase-level grounding than the reported compression baselines.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/6, 10:19 AMnot independent
    Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
  2. AggregatorHuggingFace Daily Papers7/7, 12:00 PMnot independentRepresentative
    Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval