InternVLA-A1.5 uses a native vision-language model backbone that continues training on VQA and subtask prediction, while a lightweight unified expert generates continuous actions. Future prediction is reformulated as latent querying: a small set of learnable foresight tokens extracts task-relevant future information under supervision from a frozen pretrained video-generation model. The video branch is removed at inference, avoiding pixel-level generation during control. Trained on 1.2 million robot episodes and 3 million multimodal samples, the paper reports the best overall results across six simulation benchmarks and stronger real-world compositional generalization on held-out instruction bindings, with improved long-horizon execution.
No heat snapshots are available in the last 24 hours.