Track4Action uses a frozen world-centric 3D tracker as privileged supervision for a vision-language-action policy. During training, Track4World encodes the realized geometry, motion, visibility, and camera changes in an aligned demonstration clip; learnable track queries reconstruct that representation from current VLA hidden states and condition a flow-matching action head. The tracker and future clip are removed at deployment. The supplied abstract reports 82.3% zero-shot success on LIBERO-Plus, 80.44%/81.48% on RoboTwin 2.0, and 67.5% across four physical bimanual tasks. However, arXiv:2608.03727 is dated August 2026, so the paper and claims are not currently verifiable.
No heat snapshots are available in the last 24 hours.