StatePlay argues that game world models must model explicit mechanics, not only produce visually plausible frames. The proposed model jointly predicts visual observations and internal game states such as health points, skill meters, and timers. Its mixture-of-transformers architecture keeps specialized visual and state representations while enabling cross-modal interaction, allowing predicted states to guide frame generation. According to the paper summary, StatePlay achieves an average normalized L1 state-prediction distance below 0.06 and improves mechanics fidelity in generated rollouts by 18.6% compared with models without explicit state modeling.
No heat snapshots are available in the last 24 hours.