3DZip addresses the thousands of geometry-aware tokens produced by 3D vision-language models. Its three-stage pipeline performs coarse voxelization, selects feature-diverse anchor tokens with a Determinantal Point Process, and merges remaining tokens under spatial constraints. The authors report consistent improvements over existing compression methods on three 3D question-answering benchmarks, while retaining 94.7% of the original performance. The provided abstract does not specify the benchmark names, compression ratios, latency, memory savings, or detailed ablation results.
No heat snapshots are available in the last 24 hours.