Memory Decoder at Scale extends a parametric long-term memory module to 6.9B parameters and pretrains it on 300B tokens. Because conventional Faiss indexing and retrieval becomes infeasible at this scale, the work introduces a distributed indexing and retrieval pipeline plus sparse, batch-wise loading of kNN distributions. Across 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises the average score from 29.86 to 37.34, slightly exceeding Pythia-12B at 37.24 with 39% fewer total parameters. Experiments with Qwen3 Base models also report gains from separately scaled domain memories.
No heat snapshots are available in the last 24 hours.