This paper formalizes the failure of fine-tuning to transfer newly memorized facts into downstream reasoning as the Knowing–Using Gap, covering both an accuracy gap and a temporal lag between memorization and generalization. Using a new intervention method called self-patching, the authors identify activation locations whose relocated representations improve failed cases. Their results support a knowledge-circuit misalignment hypothesis: the model may internally contain the memorized representation, but fail to route it to computation-effective layers. A simple heuristic recovers 58–75% of oracle headroom in generalization failures, with experiments conducted across domains.
No heat snapshots are available in the last 24 hours.