OmniDelta is a training-free framework for allocating a fixed retained-token budget in audio-video OmniLLMs. It first uses query intent and modality-specific skill pools to shift budget between audio and video, then assigns local budgets using content complexity and temporal redundancy. The abstract reports results on four audio-video benchmarks and two Qwen2.5-Omni models. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reportedly reduces GPU memory usage by 22.0% and delivers a 1.64x end-to-end speedup over full-token inference, while combining with existing token-pruning methods.
No heat snapshots are available in the last 24 hours.