AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment
AI Summary
This study surveyed 84 thesis supervisors across four disciplines about the relative importance of 35 assessment criteria. The resulting weights were compared with the default weights used by RubiSCoT and tested in several calibration configurations on 80 German-language theses. The best configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, but the improvement was not statistically significant. Inter-supervisor deviation was substantially lower at 4.44%, indicating that criterion-weight calibration alone is unlikely to close the alignment gap between automated and human thesis assessment.
Why it's worth reading
It provides a concrete calibration result for AI thesis assessment: supervisor-derived weights produced only a small, statistically nonsignificant improvement, while AI-human disagreement remained much larger than disagreement among supervisors.
Deep Read
What Happened
Original facts: The study surveyed 84 thesis supervisors across four academic disciplines and collected weights for 35 assessment criteria. It compared these supervisor-derived weights with RubiSCoT's default weights, then tested multiple calibration configurations on 80 German-language theses.
Core Technology
Original facts: The system uses rubric-based assessment, combining evaluations through criterion-specific weights. The researchers integrated supervisor-derived weights into several calibration settings and compared AI-generated evaluations with supervisor-assigned evaluations.
Key Evidence & Numbers
Original facts: The best configuration reduced mean relative deviation from 11.18% to 10.85%. The improvement was not statistically significant. Mean relative deviation among human supervisors was 4.44%, substantially lower than the AI-human deviation reported in the study.
Why It Matters
Analysis: The results suggest that misalignment in automated thesis assessment is not primarily a weighting problem. Even weights that better reflect supervisors' stated priorities may leave substantial differences in reasoning quality, contextual interpretation, scholarly contribution, and language-sensitive judgment.
Practical Impact
Analysis: Institutions should not expect a single expert-weighting exercise to make automated thesis grading reliable. Deployment should retain human review, measure errors by discipline and language, and inspect which individual criteria generate the largest disagreements.
Limitations & Uncertainty
Original facts: The study covers 84 supervisors and 80 German-language theses, while the abstract does not report disciplinary sample sizes, thesis selection, model versions, statistical-test details, or the complete results for each configuration. Unverified inference: Generalization to other countries, degree levels, languages, or newer models cannot be established from the abstract alone.
Original Sources
- arXiv abstract page
- System discussed: RubiSCoT; reference [1] is mentioned in the abstract, but its full bibliographic details are not provided there.