Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

First seen · 7/7/2026, 12:00 PMLatest activity · 7/7/2026, 12:00 PM

InternVLA-A1.5 uses a native vision-language model backbone that continues training on VQA and subtask prediction, while a lightweight unified expert generates continuous actions. Future prediction is reformulated as latent querying: a small set of learnable foresight tokens extracts task-relevant future information under supervision from a frozen pretrained video-generation model. The video branch is removed at inference, avoiding pixel-level generation during control. Trained on 1.2 million robot episodes and 3 million multimodal samples, the paper reports the best overall results across six simulation benchmarks and stronger real-world compositional generalization on held-out instruction bindings, with improved long-horizon execution.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/7, 12:00 PMnot independentRepresentative
    InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization