This paper introduces prompt-guided selective target sound localization and proposes SelectTSL, an end-to-end model that localizes only a user-specified sound in multi-source acoustic scenes. Its Prompt-Guided Selective Attention Module (PGSA) creates prompt-informed embeddings that guide enhancement of inter-channel phase difference (IPD) cues. The enhanced spatial features are fused with target magnitudes to jointly estimate direction of arrival (DoA) and target-source cardinality, including time-varying target counts. The authors report consistent improvements over baselines on synthetic data and real-world recordings, with robust generalization to real acoustic environments.
No heat snapshots are available in the last 24 hours.