SAM-MT extends Segment Anything 2 into an interactive, real-time framework for multi-target video segmentation. It represents individual objects with explicit queries while maintaining a shared global-context representation. Decoupled masked attention is used to reduce cross-target interference, and sparse memory supports temporal evolution. The method also includes strategies for occlusion handling and overlap prevention. According to the paper abstract, SAM-MT exceeds 36 FPS while segmenting 10 targets, with latency decoupled from the number of targets and performance intended to remain comparable to SAM2-based single-target baselines.
No heat snapshots are available in the last 24 hours.