LiLa-WAM is a lightweight world-action model for robotic manipulation that performs future-oriented reasoning in a compact latent space jointly shaped by future-state prediction and action generation. The model is designed for end-to-end training on a single 24GB GPU. It also introduces the Visual Transition Token (VTT), a language-free task representation encoding a task as a direction in visual feature space. The authors report a 90.48% success rate across 50 RoboTwin 2.0 tasks and additional evaluations on LIBERO and real-robot tasks.
No heat snapshots are available in the last 24 hours.