KVpop introduces a learned, fixed-budget KV-cache eviction policy that directly supervises keep-or-drop decisions. Its scorer is trained with a future-attention target that can be computed without materializing dense attention maps. The method also adds a delayed memory-based scorer, postponing scoring for a fixed number of steps to use near-future context. On AIME and HMMT mathematical reasoning with Qwen3-4B, KVpop reportedly retains 98% of full-attention performance at 75% KV-cache compression and 97% at 88% compression. Qwen3-8B achieves near-full teacher performance, according to the abstract.
No heat snapshots are available in the last 24 hours.