DSWorld introduces a Data Science World Model that predicts how candidate operations change a data science execution environment. Its framework combines structured state construction, cost-aware routing, lightweight real execution, and an LLM-based simulator for expensive operations. The authors build an 8K-scale transition trajectory dataset and propose Reflective World Model Optimization, an error-aware reinforcement learning strategy. Reported results indicate approximately 14× faster RL-based agent training, 3–6× faster search-based inference, and a 35.6% advantage over the strongest LLM baseline on transition prediction tasks.
No heat snapshots are available in the last 24 hours.