Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

First seen · 7/19/2026, 12:21 AMLatest activity · 7/19/2026, 12:21 AM

This paper studies whether discrete speech tokens in end-to-end speech language models preserve exploitable speaker identity information. It introduces Audio BERT (AuB), which builds speaker-sensitive representations from discrete codebooks, and SpInv, a two-stage inversion attack that reconstructs embeddings in an attacker-selected speaker-encoder space. Evaluations cover Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni on VoxCeleb under speaker-disjoint protocols. According to the abstract, only three seconds of frontend output is sufficient for SpInv to achieve cosine similarity above 0.70 in the target speaker-encoder space.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/19, 12:21 AMnot independentRepresentative
    Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models