The paper evaluates adaptive-compute latent world models on nine DeepMind Control tasks, using eight seeds and matched single-step and four-step training. It introduces rho, the shallowest-exit rollout error divided by full-depth rollout error. Six tasks benefit from depth, two show an inversion where shallow rollouts outperform full depth, and one is effectively flat. Early-exit supervision at only the first rollout step removes the inversion, suggesting a routing-related training mechanism. A dimensionality-only classifier predicts regimes on held-out tasks, while data quality and planner configuration substantially affect the observed tradeoff.
No heat snapshots are available in the last 24 hours.