Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

First seen · 8/6/2026, 01:53 AMLatest activity · 8/6/2026, 01:53 AM

OPD-V argues that modality imbalance limits on-policy self-distillation in multimodal large language models: when text dominates generation, the model underuses visual information and privileged supervision. The method builds a Positive Teacher from a zoomed-in image and a Negative Teacher from a masked image. Their logit differences define Positive Modality-Balance Logits Margins and a Modality-Balance Trust Region, which selects on-policy tokens for distillation. The abstract reports consistent reasoning gains across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods, alongside reduced training cost, but provides no exact improvement figures.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/5, 04:00 AMnot independent
    OPD-V: Visual On-Policy Self-Distillation with Modality Balance
  2. AggregatorarXiv8/6, 01:53 AMnot independentRepresentative
    OPD-V: Visual On-Policy Self-Distillation with Modality Balance