CAST uses changes in a game solver’s state value to create turn-level solver advantages for reinforcement learning with verifiable rewards (RLVR). The method addresses sparse final rewards by supplying process-level credit without requiring teacher logits. Under a soft-optimal solver assumption, maximizing solver advantage is equivalent to on-policy distillation from the solver using only scalar values. The paper reports that CAST beats all trained baselines on Sokoban, Minesweeper, and Rush Hour across in-domain and unseen-difficulty settings, and achieves the best average zero-shot performance on ALFWorld and WebShop.
No heat snapshots are available in the last 24 hours.