MiniWorld presents a lightweight, transparent, and reproducible framework for training streaming video world models from scratch. The framework uses a block-causal Video Diffusion Transformer to autoregressively predict future observations from historical observations and control signals. The paper targets a gap in current practice: many strong systems adapt pretrained video generators through complex post-training or distillation pipelines, while bidirectional pretraining can mismatch causal, streaming inference. MiniWorld is intended as an end-to-end baseline that can be trained with modest computational resources, potentially making world-model research easier to reproduce and extend. The supplied abstract is truncated, so implementation details and evaluation results require verification in the full paper.
No heat snapshots are available in the last 24 hours.