This paper argues that transition accuracy is an inadequate acceptance criterion for LLM-synthesized Code World Models used by planners. A model can achieve 100% transition accuracy on a sampling gate and at least 98% state accuracy on the planner’s search distribution, yet lose systematically because the less-than-1% of omitted dynamics are pivotal. The authors report a play cost of 0.091, with a seed-clustered 95% CI of [0.065, 0.117] over n=4,800, and propose a danger law combining play cost with a rarity-dependent gate-miss factor. They also report analogous failures in imperfect-information belief inference.
No heat snapshots are available in the last 24 hours.