This paper studies political-intent detection in Bengali memes, where noisy images, stylized embedded text, and limited language resources make multimodal classification difficult. Its framework first uses a vision-language model to extract OCR text, then encodes text and image features and combines them with token-to-region multi-head cross-attention. The authors also test a domain-specific political lexicon as a knowledge prior. On the PoliMemeDecode1 dataset, the approach reportedly reaches an approximately 0.94 Macro-F1, outperforming unimodal and feature-concatenation baselines. Interpretability analysis is reported to show grounding between textual semantics and visual evidence.
No heat snapshots are available in the last 24 hours.