Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

VideoRAE converts frozen video foundation model representations into generation-friendly video latents. It extracts multi-scale hierarchical features from encoders such as V-JEPA 2 and VideoMAEv2, then compresses them with a lightweight 1D self-attention projector. The method supports continuous latents for Diffusion Transformers and discrete tokens for autoregressive models through multi-codebook high-dimensional quantization. The abstract reports UCF-101 class-to-video gFVDs of 40 for an AR generator and 93 for a DiT generator, approximately 5x faster convergence than competing autoencoder baselines, and faster convergence than LTX-VAE in a controlled 2B text-to-video study.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders