This study surveyed 84 thesis supervisors across four disciplines about the relative importance of 35 assessment criteria. The resulting weights were compared with the default weights used by RubiSCoT and tested in several calibration configurations on 80 German-language theses. The best configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, but the improvement was not statistically significant. Inter-supervisor deviation was substantially lower at 4.44%, indicating that criterion-weight calibration alone is unlikely to close the alignment gap between automated and human thesis assessment.
No heat snapshots are available in the last 24 hours.