This paper identifies “Training-Distribution Hallucination” in world action models: when observations differ visually from training data, pixel-generative future prediction may hallucinate training-domain appearance instead of preserving the current scene. ST-WAM uses DINOv3 as a shared semantic representation for future prediction and history retrieval, while retaining Wan-VAE latents for fine-grained visual dynamics. Its Dual-Space Future Experts jointly predict VAE latents and DINO features, and Current-Anchored Intent Retrieval selects task-relevant evidence from recent semantic history. The abstract presents a diagnosis and architectural proposal, but does not provide quantitative results or full experimental details.
No heat snapshots are available in the last 24 hours.