The paper proposes an online, training-free token-pruning method for 3D question answering. Each input frame is projected into a shared voxel space using depth and camera pose, allowing spatially overlapping regions across views to be detected and redundant image tokens to be removed before they reach the language model. The method is applied to Qwen2.5-VL-7B and Qwen3-VL-8B. According to the abstract, it improves results on ScanQA, SQA3D, and OpenEQA-HM3D while reducing token usage by up to 50%.
No heat snapshots are available in the last 24 hours.