This paper proposes masked boundary modeling, a self-supervised pretraining method centered on sub-pixel boundary representations. The method dynamically discovers boundary-bearing tokens and uses them as masked targets for dense visual token learning. Scaling the approach produces LingBot-Vision, which is evaluated across diverse downstream vision tasks against DINOv3 as a strong baseline. The authors report that it enables the transition from LingBot-Depth 1.0 to 2.0 for depth completion and improves depth estimation, positioning boundary modeling as a scalable principle for structured spatial representation learning.
No heat snapshots are available in the last 24 hours.