OmniFocus is a training-free token compression method for omni-modal LLMs that independently estimates the importance of video and audio tokens based on the query. It targets two weaknesses in prior audio-visual compression approaches: unimodal guidance and the assumption that both modalities have similarly distributed information density over time. Experiments on four audio-visual benchmarks using the Qwen2.5-Omni model family report strong performance at low retention rates. On DailyOmni, Qwen2.5-Omni-7B retains 59.40 accuracy at 25% token retention and achieves up to 1.38x prefill speedup versus the full-token baseline.
No heat snapshots are available in the last 24 hours.