Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Attending to Multimodal Generation One Token at a Time

First seen · 7/8/2026, 12:00 PMLatest activity · 7/8/2026, 12:00 PM

This paper introduces One Token at a Time (OTaT), an analysis framework for tracking how multimodal large language models allocate attention to images, text, instructions, and previously generated tokens during autoregressive generation. Across two mainstream model families and four open-weight MLLMs of different sizes, the authors report recurring patterns: image attention rises when visual evidence is needed, instruction tokens are revisited during task transitions, and attention to generated history grows later in the response. Attention-blocking interventions support a functional role for these patterns. The paper also proposes a test-time intervention that redirects attention toward the relevant modality.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/8, 12:00 PMnot independentRepresentative
    Attending to Multimodal Generation One Token at a Time