Read original
arXivYue ZhangPapers87

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

SmartMage is a unified multimodal large language model for 3D scene understanding that dynamically selects among heterogeneous modalities according to query semantics, text-modality alignment, and modality quality. Its SMART module performs semantic-guided modality routing, while the MAGE module uses modality priors to adapt expert activation during multimodal reasoning. The paper reports state-of-the-art results across five 3D scene-understanding benchmarks and competitive performance on RGB-only video-understanding benchmarks. A diagnostic benchmark, ScanFacet, categorizes tasks by fine-grained semantics to examine which modality combinations are preferred for different types of questions.

Why it's worth reading

As embodied systems move beyond fixed multimodal inputs, SmartMage offers both a query-dependent routing mechanism and a diagnostic view of which modalities different 3D semantic tasks actually need.

Tags

3D理解多模态大模型模态路由具身智能混合专家场景理解arXiv