Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

First seen · 7/28/2026, 12:00 PMLatest activity · 7/28/2026, 12:00 PM

OmniVAE introduces a jointly trained audio-video VAE designed to align the latent spaces of two structurally different modalities. It combines reconstruction with a segment-level audio-video contrastive objective to capture temporal and semantic correspondence. It also distills features from pretrained modality-specific semantic encoders into each modality. According to the paper’s abstract, both objectives consistently improve latent-space learnability, leading to better downstream text-to-audio-video generation quality and more accurate cross-modal synchronization. The supplied abstract does not provide benchmark names, numerical gains, model scale, or computational costs.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/28, 12:00 PMnot independentRepresentative
    OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation