The paper proposes KV-cache grafting for frozen language models: verified knowledge is stored once as a byte-exact key-value state and restored into a fresh inference context without changing model weights. Under a pinned deterministic configuration, the authors report SHA-256-identical logits, zero KL divergence, and 100% argmax agreement across 50 samples. On AIME 2025, Gemma-4-12B reportedly improves from 80.0% to 93.3% after grafting a verified solution library. The paper also reports large reductions in decoding cost and a context expansion from 32,768 to 2,854,766 tokens.
No heat snapshots are available in the last 24 hours.