Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Visual Contrastive Self-Distillation

First seen · 7/24/2026, 12:00 PMLatest activity · 7/24/2026, 12:00 PM

The paper introduces Visual Contrastive Self-Distillation (VCSD), an on-policy self-distillation method that removes the need for an external teacher, privileged answers, visual evidence, or reasoning traces. An EMA teacher predicts the next-token distribution twice for each student-generated response prefix: once with the original image and once with a content-erased control. Their token-level log-probability difference identifies tokens whose likelihood is specifically increased by instance-level visual content. VCSD sharpens the original-image distribution within plausible support and distills it into the student. On ViRL39K, it improves seven-benchmark aggregates for Qwen3-VL 2B, 4B, and 8B from 62.27% to 67.04%, 71.30% to 73.16%, and 72.51% to 76.26%, respectively.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/24, 12:00 PMnot independentRepresentative
    Visual Contrastive Self-Distillation