The paper introduces S3M-based Phonological Activation Mapping (SPAM), which converts frame-level representations from self-supervised speech models into phonological feature activations such as voicing and nasality. Two lightweight, gradient-descent-free heads then perform phone recognition and segmentation. The authors report that the method needs less than one minute of phonetic transcriptions, generalizes to phones unseen during training, and achieves strong performance across diverse datasets. The supplied abstract does not provide detailed benchmark numbers, model configurations, or dataset names.
No heat snapshots are available in the last 24 hours.