The paper presents a zero-shot face-to-speech (F2S) framework that predicts and synthesizes a plausible speaker voice from a static facial image, removing the need for reference audio. A lightweight Face Adapter and soft-tuning of the upper face-encoder blocks align face-recognition features with the style space of a frozen StyleTTS 2 model. On held-out identities from the LRS3 audiovisual corpus, the authors report UTMOS scores of 3.7–4.0, compared with 3.61 for ground-truth speech. Retrieval is above chance, voice consistency is reported, and an English-trained adapter generates fluent Spanish speech without retraining.
No heat snapshots are available in the last 24 hours.