World in World: Exploring Video World Models with a Training-Free Inference Interface
Original title:World in World: Explore the World with World Models
Autoregressive video world models offer long-horizon exploration, yet altering viewpoints while preserving event synchrony and completing unobserved regions typically demands costly task-specific retraining. World in World introduces a training-free inference framework that translates heterogeneous control cues—including source observations, geometry projections, and retrieved rolling cache states—into clean visual states read directly by a frozen causal video model. Guided by a correspondence router and evidence-wise attention CFG, the method unified camera rerendering, subject motion transfer, and long-range revisits under a single pipeline.
Why it's worth reading
It bypasses costly end-to-end retraining by demonstrating how structured geometric evidence and attention routing can unlock precise camera control and long-horizon consistency in frozen causal video models.