Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingNewsWatching0 independent reports0

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

First seen · 7/20/2026, 12:00 PMLatest activity · 7/20/2026, 12:00 PM

Audio-Visual Flamingo (AV-Flamingo) is an open audio-visual large language model for joint reasoning over audio, images, and long-form videos. The work introduces Audio-Visual-Skills, a dataset containing approximately 7 million captioning and question-answer instances; a three-stage curriculum that progresses from short-range perception to long-horizon multi-event reasoning; and Temporal Audio-Visual Interleaved Chain-of-Thought, which grounds intermediate reasoning steps to timestamps. The authors report evaluations across more than 15 audio-visual, omni-modal, audio, and vision benchmarks, with strong results against similarly sized open models and competitive performance relative to larger models.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/20, 12:00 PMnot independentRepresentative
    Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos