The paper introduces BTHA, a backbone-transferable hierarchical adapter framework for text-guided medical image segmentation. It defines a stable feature-level interface over multi-scale visual features and text representations, allowing language guidance to transfer across convolutional and transformer-based vision encoders and different language encoders. BTHA combines hierarchical coarse-to-fine supervision with a Scale-Adaptive Gated Semantic Guidance adapter that regulates textual injection by resolution and suppresses redundant cross-modal responses through channel recalibration. The abstract reports improvements over strong baselines on four public datasets with modest computational overhead, but provides no dataset names or numerical results.
No heat snapshots are available in the last 24 hours.