This paper introduces Sparse Delta Memory (SDM), an architecture that extends Gated DeltaNet by replacing dense key-value outer-product updates with sparse reads and writes to a large explicit memory. The authors report that, under identical parameter counts and an isoFLOP constraint, increasing state-memory capacity improves in-context learning and long-context retrieval. They also initialize the memory with learned parameters, turning part of it into parametric memory, and report further gains on common-knowledge and reasoning tasks. The abstract does not provide benchmark names, numerical results, or implementation details.
No heat snapshots are available in the last 24 hours.