This paper introduces Structured All-Mask Prediction for multimodal large language model segmentation. Its STAMPlus model separates autoregressive dialogue from non-autoregressive mask prediction, generates a target list with explicit IDs and optional boxes, and jointly predicts multiple semantic or instance targets in a shared multiclass mask space. The abstract reports state-of-the-art results across referring, reasoning, open-vocabulary semantic, instance-aware, and remote-sensing small-target segmentation while preserving multimodal instruction following. For 12-category latency, repeated STAMP inference takes 13.50 seconds versus 5.16 seconds with STAMPlus. These claims are based on the supplied abstract and require verification against the full paper and benchmarks.
No heat snapshots are available in the last 24 hours.