Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

First seen · 8/6/2026, 12:54 AMLatest activity · 8/6/2026, 12:54 AM

BridgeVLA++ extends BridgeVLA with a unified spatio-temporal memory architecture that represents persistent spatial context and interaction history. It retains BridgeVLA’s strategy of projecting raw point clouds into multi-view images, predicting intermediate heatmaps, and preserving the input-output alignment of a pretrained vision-language model during 3D action learning. The paper reports strong results on spatial manipulation, state-of-the-art performance on two memory-dependent manipulation benchmarks, support for bimanual manipulation, and validation on an additional real-world robotic platform. The supplied abstract does not provide benchmark names, numerical results, dataset sizes, or hardware details.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/5, 04:00 AMnot independent
    BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
  2. AggregatorarXiv8/6, 12:54 AMnot independentRepresentative
    BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation