The paper introduces SKIP, an inference architecture for knowledge-intensive multimodal question answering that jointly conditions computation on the question, image, and estimated difficulty. It combines question-guided visual-token pruning, region-conditional sparse retrieval, bipartite sparse cross-attention, speculative knowledge verification, and an adaptive budget controller. According to the paper, SKIP matches or surpasses strong dense baselines on OK-VQA, A-OKVQA, InfoSeek, Encyclopedic-VQA, and ViQuAE, while reducing FLOPs by 3.4–6.8x and end-to-end latency by 2.7x.
No heat snapshots are available in the last 24 hours.