This paper audits KV-cache compression under matched query-aware and query-agnostic protocols, holding the model, compression ratio, instances, and decoding fixed. Across three open 7–9B models, it reports 144,300 paired RULER-8192 evaluations, 40,800 LongBench evaluations, and paired bootstrap analyses with 50,000 resamples. According to the supplied abstract, rankings change materially when the query is hidden during compression: among five methods sharing an attention backend, only KeyDiff consistently beats the best of three trivial baselines, while SnapKV trails a start-plus-recent-window baseline by 0.066 on average.
No heat snapshots are available in the last 24 hours.