This pre-registered audit compares conversational and pedagogical policies built on three tutor-model bases, using one fixed weak simulated student, deterministic leakage and independent-work detectors, and Claude Opus 4.8 as the frozen primary judge. General-purpose helpfulness failed to reliably distinguish pedagogical quality: pedagogy contrasts remained directionally consistent across judges where detected, while helpfulness rankings reversed between judges on two of three bases. In addition, answer-revealing turns were followed by less independent student work on every base. The paper argues that tutor evaluation should combine pedagogy-specific rubrics with process measures.
No heat snapshots are available in the last 24 hours.