First seen · 9/10/2026, 12:47 AMLatest activity · 9/10/2026, 12:47 AM
While self-supervised speech models excel at distinguishing words, strong discrimination often stems from acoustic word forms rather than genuine syntactic or semantic identities. By residualizing out phoneme information, this study demonstrates that later layers of HuBERT and wav2vec 2.0 genuinely encode form-independent word representations. This straightforward disentanglement also improves higher-order linguistic extraction in unsupervised word discovery tasks.
Event heat · last 24 hours
There are 7 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 20:00; latest heat is 0.