To address fragile and ungrounded step-by-step reasoning in compact models, this paper couples Qwen2.5-3B-Instruct with external symbolic verifiers and group-relative reinforcement learning with verifiable rewards (RLVR). A lightweight router directs logic problems to Z3 and physics tasks to unit-aware solvers, using verification signals for candidate scoring and revision. Across 438 test cases, the approach lifted reasoning depth and explainability from 50.68% to 72.20% while preserving answer accuracy. By separating generative narrative from deterministic verification, the framework provides a pragmatic design for reliable educational QA.
No heat snapshots are available in the last 24 hours.