This paper introduces LingBot-VLA 2.0, an upgraded vision-language-action model designed to reduce the gap between laboratory benchmarks and real-world robot deployment. The system uses a revised data-processing pipeline and approximately 60,000 hours of pretraining data, including 50,000 hours of robot trajectories across 20 configurations and 10,000 hours of egocentric human videos. It expands control beyond dual-arm platforms to heads, waists, mobile bases, and dexterous hands, and adds predictive dynamics modeling using video semantic priors and depth-based geometric cues. The authors report gains on the GM-100 benchmark and cross-embodiment, long-horizon mobile manipulation across two robotic platforms.
No heat snapshots are available in the last 24 hours.