Read original
arxivpapers82

AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

AI Summary

This study surveyed 84 thesis supervisors across four disciplines about the relative importance of 35 assessment criteria. The resulting weights were compared with the default weights used by RubiSCoT and tested in several calibration configurations on 80 German-language theses. The best configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, but the improvement was not statistically significant. Inter-supervisor deviation was substantially lower at 4.44%, indicating that criterion-weight calibration alone is unlikely to close the alignment gap between automated and human thesis assessment.

Why it's worth reading

It provides a concrete calibration result for AI thesis assessment: supervisor-derived weights produced only a small, statistically nonsignificant improvement, while AI-human disagreement remained much larger than disagreement among supervisors.

Deep Read

What Happened

Original facts: The study surveyed 84 thesis supervisors across four academic disciplines and collected weights for 35 assessment criteria. It compared these supervisor-derived weights with RubiSCoT's default weights, then tested multiple calibration configurations on 80 German-language theses.

Core Technology

Original facts: The system uses rubric-based assessment, combining evaluations through criterion-specific weights. The researchers integrated supervisor-derived weights into several calibration settings and compared AI-generated evaluations with supervisor-assigned evaluations.

Key Evidence & Numbers

Original facts: The best configuration reduced mean relative deviation from 11.18% to 10.85%. The improvement was not statistically significant. Mean relative deviation among human supervisors was 4.44%, substantially lower than the AI-human deviation reported in the study.

Why It Matters

Analysis: The results suggest that misalignment in automated thesis assessment is not primarily a weighting problem. Even weights that better reflect supervisors' stated priorities may leave substantial differences in reasoning quality, contextual interpretation, scholarly contribution, and language-sensitive judgment.

Practical Impact

Analysis: Institutions should not expect a single expert-weighting exercise to make automated thesis grading reliable. Deployment should retain human review, measure errors by discipline and language, and inspect which individual criteria generate the largest disagreements.

Limitations & Uncertainty

Original facts: The study covers 84 supervisors and 80 German-language theses, while the abstract does not report disciplinary sample sizes, thesis selection, model versions, statistical-test details, or the complete results for each configuration. Unverified inference: Generalization to other countries, degree levels, languages, or newer models cannot be established from the abstract alone.

Original Sources

  • arXiv abstract page
  • System discussed: RubiSCoT; reference [1] is mentioned in the abstract, but its full bibliographic details are not provided there.

Tags

论文评估教育AI自动评分Rub­iSCoT校准人机一致性高等教育