Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

First seen · 7/29/2026, 06:33 PMLatest activity · 7/29/2026, 06:33 PM

The paper presents a zero-shot face-to-speech (F2S) framework that predicts and synthesizes a plausible speaker voice from a static facial image, removing the need for reference audio. A lightweight Face Adapter and soft-tuning of the upper face-encoder blocks align face-recognition features with the style space of a frozen StyleTTS 2 model. On held-out identities from the LRS3 audiovisual corpus, the authors report UTMOS scores of 3.7–4.0, compared with 3.61 for ground-truth speech. Retrieval is above chance, voice consistency is reported, and an English-trained adapter generates fluent Spanish speech without retraining.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/29, 06:33 PMnot independentRepresentative
    Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model