This paper introduces One Token at a Time (OTaT), an analysis framework for tracking how multimodal large language models allocate attention to images, text, instructions, and previously generated tokens during autoregressive generation. Across two mainstream model families and four open-weight MLLMs of different sizes, the authors report recurring patterns: image attention rises when visual evidence is needed, instruction tokens are revisited during task transitions, and attention to generated history grows later in the response. Attention-blocking interventions support a functional role for these patterns. The paper also proposes a test-time intervention that redirects attention toward the relevant modality.
No heat snapshots are available in the last 24 hours.