This paper proposes Task-Agnostic Pretraining (TAP), which separates VLA learning into physical competence and language grounding. TAP first learns transferable motor priors from unlabeled interaction data, including discarded off-task trajectories and autonomous play, using a self-supervised inverse-dynamics objective. A lightweight second stage then aligns those priors with language using limited expert demonstrations. According to the supplied abstract, TAP matches models trained on more than 1 million expert trajectories on SIMPLER, improves standard behavior cloning by 10 percentage points, and achieves 25% success on a real WidowX under camera perturbations, while internet-scale baselines fall to 0%.
No heat snapshots are available in the last 24 hours.