RestoreKV augments query-agnostic KV cache eviction with a learned, context-conditioned restoration cache. After context prefill, a small set of restore tokens attends to the full KV cache in one LoRA-adapted pass, while the original scorer and eviction rule remain unchanged. The adapters are disabled during later queries and decoding. Training updates only 0.4% of parameters through parameter-efficient self-distillation from a frozen full-cache model. Across four backbones and four long-context benchmarks, RestoreKV improved 59 of 60 paired budget-matched settings on Qwen3-4B. At a 5% budget, KVzip improved from 38.2 to 73.2 RULER-4K accuracy.
No heat snapshots are available in the last 24 hours.