SmartMage is a unified multimodal large language model for 3D scene understanding that dynamically selects among heterogeneous modalities according to query semantics, text-modality alignment, and modality quality. Its SMART module performs semantic-guided modality routing, while the MAGE module uses modality priors to adapt expert activation during multimodal reasoning. The paper reports state-of-the-art results across five 3D scene-understanding benchmarks and competitive performance on RGB-only video-understanding benchmarks. A diagnostic benchmark, ScanFacet, categorizes tasks by fine-grained semantics to examine which modality combinations are preferred for different types of questions.
No heat snapshots are available in the last 24 hours.