WrAFT is an automated writing evaluation system that separates argumentative-essay assessment into scoring, surface-level feedback, and deep-level feedback modules. The authors evaluate LLaMA-3.3-70B-Instruct, GPT-4o, and Claude 3.7 using direct prompting and supervised fine-tuning. On a proprietary dataset of 480 TOEFL Independent Writing essays scored on a 0–5 scale, WrAFT reportedly achieves a quadratic weighted kappa of 0.84 and an RMSE of 0.44 against official scores. Human reviewers approved 96.14% of surface feedback, 93.03% of macro feedback, and 94.69% of micro feedback. A free interactive interface is also available.
No heat snapshots are available in the last 24 hours.