The WanSong v1.0 technical report presents a pure-diffusion foundation model for long-form music generation. It claims direct generation of high-fidelity multilingual songs up to five minutes, with vocal and background-music stems produced in a single run. The framework also uses step distillation for faster inference and is designed to support fine-tuning, customization, and downstream editing. The supplied abstract does not provide model size, datasets, benchmark results, licensing details, or evidence comparing WanSong with existing systems.
No heat snapshots are available in the last 24 hours.