BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
BridgeVLA++ extends BridgeVLA with a unified spatio-temporal memory architecture that represents persistent spatial context and interaction history. It retains BridgeVLA’s strategy of projecting raw point clouds into multi-view images, predicting intermediate heatmaps, and preserving the input-output alignment of a pretrained vision-language model during 3D action learning. The paper reports strong results on spatial manipulation, state-of-the-art performance on two memory-dependent manipulation benchmarks, support for bimanual manipulation, and validation on an additional real-world robotic platform. The supplied abstract does not provide benchmark names, numerical results, dataset sizes, or hardware details.
Why it's worth reading
3D VLA systems are moving beyond single-step perception toward history-dependent manipulation. This work is timely because it combines explicit spatio-temporal memory with data efficiency, distribution-shift robustness, bimanual tasks, and validation on another physical robot platform.