Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

First seen · 7/29/2026, 09:33 PMLatest activity · 7/29/2026, 09:33 PM

The paper introduces Counterfactual Modality Attribution (CMA), a framework for estimating whether images, text, or their combination drives an MLLM prediction. It creates image-only, text-only, and joint multimodal counterfactuals with coupled diffusion priors, then derives modality contributions using Shapley values from cooperative game theory. On controlled synthetic benchmarks with known modality reliance and a real-world multimodal clinical dataset, the authors report 98% accuracy in identifying the decision-driving modality in controlled cases and consistent improvements over baselines. The work positions modality attribution as complementary to token- or region-level explanations, especially for auditing shortcut reasoning in safety-critical systems.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/29, 09:33 PMnot independentRepresentative
    Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs