BridgeVLA++ extends BridgeVLA with a unified spatio-temporal memory architecture that represents persistent spatial context and interaction history. It retains BridgeVLA’s strategy of projecting raw point clouds into multi-view images, predicting intermediate heatmaps, and preserving the input-output alignment of a pretrained vision-language model during 3D action learning. The paper reports strong results on spatial manipulation, state-of-the-art performance on two memory-dependent manipulation benchmarks, support for bimanual manipulation, and validation on an additional real-world robotic platform. The supplied abstract does not provide benchmark names, numerical results, dataset sizes, or hardware details.
No heat snapshots are available in the last 24 hours.