This paper argues that behavior-cloning finetuning can overwrite the visual and semantic representations inherited from a pretrained vision-language model, while web image-text co-training leaves language and action objectives misaligned because they use separate observations. Anchor-Align adds two objectives: Vision-Language Anchoring distills layer-wise representations from a frozen VLM copy, and Language-Action Alignment converts action targets into discrete motion-direction labels so language and action are trained on the same robot observation. On a physical xArm7 robot, the abstract reports success-rate gains from 28% to 54% and from 37% to 60% across two VLA architectures, with additional simulated gains on LIBERO-PRO, LIBERO-Plus, and CALVIN.
No heat snapshots are available in the last 24 hours.