Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

First seen · 7/15/2026, 05:37 PMLatest activity · 7/15/2026, 05:37 PM

The paper introduces GMoT, a gated motion-aware tokenization module for micro-gesture video reasoning with multimodal LLMs. It combines spatially weighted pooling, adjacent-frame temporal differencing, and a conservatively initialized semantic gate to compress sparse kinematic evidence before temporal modeling. The method is paired with semi-supervised anatomically focused captions and progressive reward-guided policy refinement. On iMiGUE and SMG, it reports Top-1 accuracies of 67.32% and 73.11%, improving a Qwen3-VL-8B baseline by 6.80 and 3.11 percentage points, respectively. It also proposes BRG Recall and an overlapping-label cross-domain transfer protocol.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/15, 05:37 PMnot independentRepresentative
    GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs