This paper presents a unified guidance framework for flow-matching speech synthesis. Its data-guidance component uses heterogeneous augmentation to encourage disentanglement between linguistic content and acoustic residue. Its model-guidance component combines trajectory rectification with a new intrinsic guidance objective, distilling conditional information into network weights while straightening the inference path. According to the abstract, the method removes the inference overhead of Classifier-Free Guidance (CFG), accelerates inference by nearly three times, and improves speaker similarity against state-of-the-art baselines.
No heat snapshots are available in the last 24 hours.