The paper introduces WorldWeaver (W²), a streaming multi-agent video diffusion model that maintains cross-agent world state registers during rollout. These learnable tokens store shared world information, individual agent status, and evolving scene state, and are updated after each generated video chunk. The model also uses a Mixture-of-Transformers architecture with separate parameters for world-state modeling and visual-frame modeling. Supervision covers individual status, global views including bird’s-eye views, and scene text. Experiments on two-agent Minecraft video generation report improved logical consistency and generation quality, although the abstract does not provide numerical results or comparisons.
No heat snapshots are available in the last 24 hours.