This paper studies XAI-guided adaptive fusion (XGAF), a tree-based mixture of unimodal and cross-modal experts whose sample-level weights come from TreeSHAP attribution magnitudes. It finds that mean-absolute and median-absolute reductions can underweight high-dimensional cross-modal experts, while sum-absolute reduction preserves their total attribution mass. On MELD, sum-abs XGAF reaches 0.5983 weighted F1 with a Transformer aggregator, close to early fusion at 0.6018 and above probability-average late fusion at 0.4598. On CMU-MOSEI, it reaches 0.6519 versus 0.6485 and 0.5696. Ablations attribute most gains to adding cross-modal experts, particularly the trimodal expert.
No heat snapshots are available in the last 24 hours.