This paper studies how fine-tuning Vision-Language-Action (VLA) models on limited robot demonstrations can erode the semantic structure inherited from pretrained Vision-Language Models. It proposes a plug-and-play training method that anchors action representations to a semantic manifold and separates them into shared semantic and private channels. These auxiliary components are discarded at inference, so the deployed model remains unchanged. Across multiple VLA backbones, simulation tasks, and real-world benchmarks, the authors report improvements of up to 18.7% on real-world in-distribution tasks and 21.5% on out-of-distribution generalization.
No heat snapshots are available in the last 24 hours.