Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

First seen · 7/26/2026, 02:36 PMLatest activity · 7/26/2026, 02:36 PM

This paper studies political-intent detection in Bengali memes, where noisy images, stylized embedded text, and limited language resources make multimodal classification difficult. Its framework first uses a vision-language model to extract OCR text, then encodes text and image features and combines them with token-to-region multi-head cross-attention. The authors also test a domain-specific political lexicon as a knowledge prior. On the PoliMemeDecode1 dataset, the approach reportedly reaches an approximately 0.94 Macro-F1, outperforming unimodal and feature-concatenation baselines. Interpretability analysis is reported to show grounding between textual semantics and visual evidence.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/26, 02:36 PMnot independentRepresentative
    Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation