MEPA targets representation-learning conflicts in visual autoregressive models, where low-resolution scales capture semantics, high-resolution scales model details, and early semantic errors propagate through causal generation. It introduces a scale-aware, token-routed Mixture-of-Experts architecture for scale-adaptive expert selection, together with external self-supervised features for stronger early-scale semantics. The method uses residual feature aggregation rather than naive feature alignment. The authors report better FID than a dense baseline on ImageNet 256×256, with roughly half the default training epochs and a smaller parameter budget, while adding only marginal training cost. The abstract does not provide exact metrics.
No heat snapshots are available in the last 24 hours.