GigaWorld-Policy-0.5 presents an action-centered World Action Model (WAM) for robot control. It uses future visual dynamics as training supervision but decodes actions only at inference, avoiding the cost of explicitly generating future videos. The system combines mixed Action-Conditioned World Modeling (AC-WM) and WAM pretraining with a Mixture-of-Transformers architecture that assigns visual-dynamics modeling and action generation to specialized experts. The paper reports 85 ms inference latency on a local RTX 4090 and introduces an agent-based AutoResearch pipeline for searching training configurations.
No heat snapshots are available in the last 24 hours.