Read original
HuggingFace Daily PapersBohai GuPapers86

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

WorldCycle introduces a self-verifiable reinforcement learning framework for long-horizon video world models. It uses reversible action cycles: an action sequence followed by its inverse should analytically return the system to its initial state, creating annotation-free supervision for long-horizon consistency. The method combines a spatial closure reward, which compares mirrored forward and reverse segments, with a temporal consistency reward across repeated cycle executions. The authors also release CycleBench for evaluating state-returning ability under complex action structures. The abstract reports up to a 44% reduction in state-returning drift and nearly 4x composite-action accuracy over the base model.

Why it's worth reading

Long-horizon world models lack reliable future-state supervision; WorldCycle turns reversible actions into annotation-free rewards and directly targets that verification bottleneck.

Tags

世界模型强化学习视频生成长时程规划可验证学习CycleBench