Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

First seen · 8/4/2026, 12:00 PMLatest activity · 8/4/2026, 12:00 PM

SwanTale is a multi-speaker expressive speech and audio generation model designed for both instruction-based and zero-shot tasks. The instruction setting accepts descriptions of environments, speaker styles, and fine-grained content, while zero-shot generation additionally uses reference audio. The authors introduce SwanData-Caption for cleaned, synthetically expanded, and hierarchically captioned data, plus SwanVAE, reward-conditioned quality control, Engram conditioning, Unified MoE, curriculum learning, and GRPO post-training. The abstract reports leading results across multiple instruct and zero-shot metrics, including the best expressiveness scores, but does not provide numerical results or detailed baseline information.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers8/4, 12:00 PMnot independentRepresentative
    SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks