Valdi, or Value Diffusion World Models, combines end-to-end online training for model predictive control with a latent diffusion dynamics model. The method targets the tension between expressive uncertainty modeling and the low latency required for online planning. In preliminary CarRacing experiments, Valdi uses a single diffusion step during both training and inference and matches a deterministic MLP baseline. The results also reveal a trade-off between predictive multimodality and control performance. The authors provide code publicly, but the evaluation remains limited to the reported preliminary setting.
No heat snapshots are available in the last 24 hours.