This paper examines whether speaker anonymization remains effective when an attacker can aggregate multiple utterances and modalities. Audio-only speaker verification improves consistently as more anonymized speech becomes available. Adding prosodic and linguistic information further improves performance over unimodal systems, while frame-level aggregation produces the lowest equal error rates (EERs) among the compared strategies. With only five anonymized utterances, combining audio and text reduces EER by more than 15% relative to audio-only aggregation. The findings indicate that anonymization focused mainly on vocal characteristics may leave speaker-discriminative information accessible through repeated interactions and content-related cues.
No heat snapshots are available in the last 24 hours.