The paper introduces a step-level reward framework for physics reasoning in small language models. During training, an external verifier identifies the first reasoning error, produces targeted structured feedback, and guides revision with policy gradients and KL regularization. The method does not use ground-truth solutions as generation targets and does not require preference-data construction; the verifier is used only during training. The abstract reports 17–20% gains over chain-of-thought prompting across five physics benchmarks, a reduction in calculation errors from 56.9% to 23.5%, and a reduction in miscomprehension errors from 22.3% to 12.0% in the best cases. Conceptual errors remain difficult, decreasing from 89.7% to 68.7%.
No heat snapshots are available in the last 24 hours.