This paper addresses revisit inconsistency in long-horizon autoregressive video rendering. When a bounded KV cache evicts earlier context, revisiting a previously seen location can produce a different appearance even when depth or geometry conditioning is unchanged. The proposed training-free method uses pose-matched temporal correspondence to retrieve historical latent chunks into the KV cache, then uses camera pose and depth reprojection to bias token attention toward geometrically corresponding regions. Experiments on loop-closure trajectories from TartanAir and TartanGround reportedly outperform existing training-free baselines on revisit consistency without reducing overall video quality.
No heat snapshots are available in the last 24 hours.