Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models

First seen · 7/15/2026, 04:54 PMLatest activity · 7/15/2026, 04:54 PM

This paper argues that transition accuracy is an inadequate acceptance criterion for LLM-synthesized Code World Models used by planners. A model can achieve 100% transition accuracy on a sampling gate and at least 98% state accuracy on the planner’s search distribution, yet lose systematically because the less-than-1% of omitted dynamics are pivotal. The authors report a play cost of 0.091, with a seed-clustered 95% CI of [0.065, 0.117] over n=4,800, and propose a danger law combining play cost with a rarity-dependent gate-miss factor. They also report analogous failures in imperfect-information belief inference.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/15, 04:54 PMnot independentRepresentative
    When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models