This paper presents a hierarchical self-supervised world model for symbolic music, built around a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives. Frozen representation probes reportedly align model depth with musical time scale: coarse levels expose phrase boundaries, while fine levels encode density and harmonic detail. A small chord-supervision head improves joint chord recovery from 0.18 to 0.54 and unsupervised key detection from 0.16 to 0.70. A conditional flow-matching generator supports reconstruction, controlled variation, and masked graphical inpainting, with reported latency of 2.8 seconds on CPU and 0.6 seconds on Apple MPS.
No heat snapshots are available in the last 24 hours.