Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

First seen · 7/8/2026, 12:00 PMLatest activity · 7/8/2026, 12:00 PM

This paper proposes a parallelized autoregressive framework for omni-modal dense video captioning. Its central observation is that temporally distinct events often have weak local dependencies, allowing tokens across events to be decoded in parallel while preserving sequential decoding within each event. The method introduces latent global planning to learn event-level structure and compact inter-event causal representations, followed by event-factorized parallel decoding with local and global awareness. The abstract reports improvements in efficiency and performance across multiple benchmarks, but does not provide acceleration ratios, metric values, model sizes, or benchmark names.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/3, 01:13 PMnot independent
    Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
  2. AggregatorHuggingFace Daily Papers7/8, 12:00 PMnot independentRepresentative
    Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning