This paper studies whether discrete speech tokens in end-to-end speech language models preserve exploitable speaker identity information. It introduces Audio BERT (AuB), which builds speaker-sensitive representations from discrete codebooks, and SpInv, a two-stage inversion attack that reconstructs embeddings in an attacker-selected speaker-encoder space. Evaluations cover Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni on VoxCeleb under speaker-disjoint protocols. According to the abstract, only three seconds of frontend output is sufficient for SpInv to achieve cosine similarity above 0.70 in the target speaker-encoder space.
No heat snapshots are available in the last 24 hours.