The paper introduces GMoT, a gated motion-aware tokenization module for micro-gesture video reasoning with multimodal LLMs. It combines spatially weighted pooling, adjacent-frame temporal differencing, and a conservatively initialized semantic gate to compress sparse kinematic evidence before temporal modeling. The method is paired with semi-supervised anatomically focused captions and progressive reward-guided policy refinement. On iMiGUE and SMG, it reports Top-1 accuracies of 67.32% and 73.11%, improving a Qwen3-VL-8B baseline by 6.80 and 3.11 percentage points, respectively. It also proposes BRG Recall and an overlapping-label cross-domain transfer protocol.
No heat snapshots are available in the last 24 hours.