DynaVieW introduces a dynamic schema-guided world model for predicting and simulating visual changes across videos or multi-image sequences. It represents keyframes as broad scene states and organizes transitions into a hierarchical schema covering actions and resulting environmental changes. The model jointly learns transition prediction and state simulation with a mixture-of-experts architecture, cross-expert selective attention, and schema-token reweighted loss. According to the provided abstract, DynaVieW improves consistency, controllability, and instruction following in visual narrative generation and world simulation, although quantitative results are not included in the supplied summary.
No heat snapshots are available in the last 24 hours.