This report introduces a unified framework for full-length music generation across lyrics-to-song, instrumental generation, and cover-song transformation. Its architecture combines a semantic-aware tokenizer, an eight-codebook RVQ representation, hierarchical autoregressive modeling through “hybird-LM,” continuous full-song rendering with FullDiT in VAE latent space, and a two-level melody module for cover generation. The authors also investigate DPO, GRPO, OPD, and flow-based GRPO as post-training strategies. Evaluation uses a multilingual automatic benchmark and the Artificial Analysis Music with Vocals leaderboard, where the system reportedly achieves competitive results.
No heat snapshots are available in the last 24 hours.