SwanTale is a multi-speaker expressive speech and audio generation model designed for both instruction-based and zero-shot tasks. The instruction setting accepts descriptions of environments, speaker styles, and fine-grained content, while zero-shot generation additionally uses reference audio. The authors introduce SwanData-Caption for cleaned, synthetically expanded, and hierarchically captioned data, plus SwanVAE, reward-conditioned quality control, Engram conditioning, Unified MoE, curriculum learning, and GRPO post-training. The abstract reports leading results across multiple instruct and zero-shot metrics, including the best expressiveness scores, but does not provide numerical results or detailed baseline information.
No heat snapshots are available in the last 24 hours.