Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

ViPo-MLLM: A Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

First seen · 7/4/2026, 09:17 AMLatest activity · 7/4/2026, 09:17 AM

ViPo-MLLM targets gloss-free sign language translation by combining spatio-temporal RGB video with human-pose features. Dedicated encoders model intra-modal dynamics, while cross-modal attention captures longer-range relationships between visual and pose signals. A structured prompt conditions the fused representation before an LLM generates spoken-language sentences, trained with contrastive and language-modeling objectives. The paper reports new state-of-the-art results on PHOENIX14T and CSL-Daily, with competitive performance against gloss-based recognition approaches. Exact scores are not included in the supplied abstract.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/4, 09:17 AMnot independentRepresentative
    ViPo-MLLM: A Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation