The paper introduces MultAttnAttrib, a training-free method for attributing long-document QA answers to multimodal evidence. It uses the model’s prefill pass, selected attention heads, and calibrated thresholds to locate source regions without an additional attribution model or training stage. The authors also introduce MultAttrEval, a benchmark with fine-grained ground-truth annotations for answer components grounded in multimodal documents. According to the abstract, MultAttnAttrib outperforms several prompting-based attribution methods, approaches the reported performance of GPT 5.4, and reduces direct-inference latency to as little as one-seventh of prompting on the same base model.
No heat snapshots are available in the last 24 hours.