SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
SmartMage is a unified multimodal large language model for 3D scene understanding that dynamically selects among heterogeneous modalities according to query semantics, text-modality alignment, and modality quality. Its SMART module performs semantic-guided modality routing, while the MAGE module uses modality priors to adapt expert activation during multimodal reasoning. The paper reports state-of-the-art results across five 3D scene-understanding benchmarks and competitive performance on RGB-only video-understanding benchmarks. A diagnostic benchmark, ScanFacet, categorizes tasks by fine-grained semantics to examine which modality combinations are preferred for different types of questions.
Why it's worth reading
As embodied systems move beyond fixed multimodal inputs, SmartMage offers both a query-dependent routing mechanism and a diagnostic view of which modalities different 3D semantic tasks actually need.