This paper proposes a speech-forensics method based on vowel-spectrum distributions. Using Japanese, whose moraic syllabary provides a relatively constrained five-vowel system, it models normalized speech spectra as probability density functions over cochlear frequency bands. The method measures distances between vowels with the Wasserstein metric, then applies topological mapping and persistent homology. The paper argues that generative-AI speech has shorter inter-vowel Wasserstein distances because synthesis is constrained by a limited set of training spectra, whereas natural speech shows greater articulatory and spectral diversity. Part 1 presents the method and examples, but the supplied abstract does not report benchmark accuracy, dataset scale, or comparisons with existing detectors.
No heat snapshots are available in the last 24 hours.