The paper reports that VGGT, despite receiving no explicit co-visibility supervision, internally develops representations useful for deciding whether image pairs share visible surfaces. Early layers appear to build 3D-aware scene representations, while later layers perform co-visibility reasoning; layer L17 consistently acts as a negative anchor for non-co-visible pairs. The authors introduce Co-VGGT, a lightweight layer-wise mixture-of-experts head with fewer than 7.5M trainable parameters while freezing VGGT. On Co-VisiON, it reportedly exceeds the human annotation baseline and prior work, with pairwise ECE of 0.030.
No heat snapshots are available in the last 24 hours.