This paper isolates music tokenization by holding pretrained Qwen3.5 backbones, data, training budget, and decoding fixed across seven representations. The authors report that representation dominates model size for distributional fidelity: a 0.8B model using the released Performance Music Tokens (PMT) reaches FMD 159, versus 272–286 for beat-grid tokenizations, and outperforms a 27B beat-grid model. PMT uses 10 ms timing, per-note velocity, multi-track texture, and a 609-symbol vocabulary. The work also reports replication on a 26M from-scratch backbone and another performance-resolution tokenizer, while explicitly leaving audible superiority to a preregistered human study.
No heat snapshots are available in the last 24 hours.