GEAR jointly trains a vector-quantized tokenizer and an autoregressive image generator instead of freezing the tokenizer before generator training. Its dual read-out uses a hard one-hot branch for next-token prediction and a differentiable soft branch for representation alignment, allowing gradients to guide the tokenizer without backpropagating through discrete indices. The paper reports up to 10x faster ImageNet gFID convergence than the LlamaGen-REPA baseline, stronger patch-level and spatially coherent features, generalization across VQVAE, LFQ, and IBQ quantizers, and extension to text-to-image generation.
No heat snapshots are available in the last 24 hours.