The paper introduces WMBench, a benchmark built from real-robot teleoperation data and matched policy rollouts for evaluating world models as surrogate robot-policy evaluators. It compares 7 video world models, 4 action representations, rollout horizons, and evaluation metrics across more than 324,000 simulated rollouts paired with real executions. The authors report that long-horizon, action-faithful consistency matters more than short-term visual realism. They also study pretraining, memory, action encoding, and evaluator-focused post-training, then present GigaWorld-1 and release code, models, datasets, and toolkits.
No heat snapshots are available in the last 24 hours.