AlayaWorld is an interactive long-horizon video world model built on a 15B video diffusion Transformer. It autoregressively generates 24-fps video at 540p and 720p in short latent chunks, conditioned on camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. The authors train with corrupted histories and rollout prediction residuals to reduce long-term drift. A discrete autoregressive distillation scheme combining distribution matching, self-forcing++, and consistency distillation reduces inference from roughly 30 sampling steps to four per chunk. The report claims best performance on iWorld-Bench.
No heat snapshots are available in the last 24 hours.