This paper studies whether cropping student responses with bounding boxes improves small language model performance on vision-based grading. Using scanned handwritten responses from the 2025 Australian Physics Olympiad, the authors evaluate models ranging from 4B to 72B parameters under different chain-of-thought prompting and image-cropping settings. The abstract reports that bounding-box cropping improves grading accuracy and reduces computational cost, measured in FLOPs, across the evaluated models. The work frames answer-region extraction as a practical preprocessing step for privacy-conscious, scalable educational assessment with smaller multimodal systems.
No heat snapshots are available in the last 24 hours.