VideoRAE converts frozen video foundation model representations into generation-friendly video latents. It extracts multi-scale hierarchical features from encoders such as V-JEPA 2 and VideoMAEv2, then compresses them with a lightweight 1D self-attention projector. The method supports continuous latents for Diffusion Transformers and discrete tokens for autoregressive models through multi-codebook high-dimensional quantization. The abstract reports UCF-101 class-to-video gFVDs of 40 for an AR generator and 93 for a DiT generator, approximately 5x faster convergence than competing autoencoder baselines, and faster convergence than LTX-VAE in a controlled 2B text-to-video study.
No heat snapshots are available in the last 24 hours.