Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

First seen · 7/29/2026, 12:00 PMLatest activity · 7/29/2026, 12:00 PM

OmniDelta is a training-free framework for allocating a fixed retained-token budget in audio-video OmniLLMs. It first uses query intent and modality-specific skill pools to shift budget between audio and video, then assigns local budgets using content complexity and temporal redundancy. The abstract reports results on four audio-video benchmarks and two Qwen2.5-Omni models. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reportedly reduces GPU memory usage by 22.0% and delivers a 1.64x end-to-end speedup over full-token inference, while combining with existing token-pruning methods.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/29, 12:00 PMnot independentRepresentative
    OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs