This paper proposes a parallelized autoregressive framework for omni-modal dense video captioning. Its central observation is that temporally distinct events often have weak local dependencies, allowing tokens across events to be decoded in parallel while preserving sequential decoding within each event. The method introduces latent global planning to learn event-level structure and compact inter-event causal representations, followed by event-factorized parallel decoding with local and global awareness. The abstract reports improvements in efficiency and performance across multiple benchmarks, but does not provide acceleration ratios, metric values, model sizes, or benchmark names.
No heat snapshots are available in the last 24 hours.