This pilot study presents a privacy-aware and computationally efficient framework for recognizing classroom incidents from CCTV-style observations. It introduces a hybrid benchmark combining generative CCTV-style videos with real-world classroom pose data. The method builds hierarchical kinematic representations focused on motion direction, speed, acceleration, and intensity, then distills multi-order motion reasoning from a large teacher model into a lightweight single-order student. According to the abstract, the student outperforms substantially larger baselines at less than one-tenth of their computational cost, with stronger out-of-domain reasoning and zero-shot synthetic-to-real generalization. The authors state that the benchmark, code, and tools will be released publicly.
No heat snapshots are available in the last 24 hours.