Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal

First seen · 7/30/2026, 08:37 PMLatest activity · 7/30/2026, 08:37 PM

This pre-registered audit compares conversational and pedagogical policies built on three tutor-model bases, using one fixed weak simulated student, deterministic leakage and independent-work detectors, and Claude Opus 4.8 as the frozen primary judge. General-purpose helpfulness failed to reliably distinguish pedagogical quality: pedagogy contrasts remained directionally consistent across judges where detected, while helpfulness rankings reversed between judges on two of three bases. In addition, answer-revealing turns were followed by less independent student work on every base. The paper argues that tutor evaluation should combine pedagogy-specific rubrics with process measures.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/30, 08:37 PMnot independentRepresentative
    Rethinking LLM-Judged Helpfulness as a Pedagogy Signal