Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models

First seen · 7/3/2026, 03:41 PMLatest activity · 7/3/2026, 03:41 PM

OmniFocus is a training-free token compression method for omni-modal LLMs that independently estimates the importance of video and audio tokens based on the query. It targets two weaknesses in prior audio-visual compression approaches: unimodal guidance and the assumption that both modalities have similarly distributed information density over time. Experiments on four audio-visual benchmarks using the Qwen2.5-Omni model family report strong performance at low retention rates. On DailyOmni, Qwen2.5-Omni-7B retains 59.40 accuracy at 25% token retention and achieves up to 1.38x prefill speedup versus the full-token baseline.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/3, 03:41 PMnot independentRepresentative
    OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models