The Qwen-Music technical report presents a music-generation system for text-to-music and cover-song generation with complete vocal singing. Its architecture combines a Qwen-Music-Tokenizer, a Qwen-Music-LLM, and a Qwen-Music-Render module. Audio is compressed into a 25 Hz single-codebook stream of semantic tokens, while Melody-CoT plans melody tokens before full-song generation. Training used more than 5 million hours of multilingual music spanning hundreds of languages, followed by supervised initialization, offline DPO, and online GSPO. On 600 Chinese and English prompts, the report says Qwen-Music led in 13 of 16 objective musicality and audio-quality metrics, with professional evaluators preferring it over leading proprietary systems.
No heat snapshots are available in the last 24 hours.