arXivRobin Huo
Do Speech Foundation Models Really Learn Words?
Original title:Do speech foundation models really learn words?
Papers74
While self-supervised speech models excel at distinguishing words, strong discrimination often stems from acoustic word forms rather than genuine syntactic or semantic identities. By residualizing out phoneme information, this study demonstrates that later layers of HuBERT and wav2vec 2.0 genuinely encode form-independent word representations. This straightforward disentanglement also improves higher-order linguistic extraction in unsupervised word discovery tasks.
Why it's worth reading
It answers a foundational interpretability question in speech AI by proving that foundation models build genuine word-level abstractions distinct from surface phonetic forms.
Tags
语音基础模型HuBERTwav2vec 2.0表征学习可解释性计算语言学