ViPo-MLLM targets gloss-free sign language translation by combining spatio-temporal RGB video with human-pose features. Dedicated encoders model intra-modal dynamics, while cross-modal attention captures longer-range relationships between visual and pose signals. A structured prompt conditions the fused representation before an LLM generates spoken-language sentences, trained with contrastive and language-modeling objectives. The paper reports new state-of-the-art results on PHOENIX14T and CSL-Daily, with competitive performance against gloss-based recognition approaches. Exact scores are not included in the supplied abstract.
No heat snapshots are available in the last 24 hours.