WorldCycle introduces a self-verifiable reinforcement learning framework for long-horizon video world models. It uses reversible action cycles: an action sequence followed by its inverse should analytically return the system to its initial state, creating annotation-free supervision for long-horizon consistency. The method combines a spatial closure reward, which compares mirrored forward and reverse segments, with a temporal consistency reward across repeated cycle executions. The authors also release CycleBench for evaluating state-returning ability under complex action structures. The abstract reports up to a 44% reduction in state-returning drift and nearly 4x composite-action accuracy over the base model.
No heat snapshots are available in the last 24 hours.