This survey proposes an architecture-centric taxonomy for memory in large language models. It organizes existing approaches across three orthogonal dimensions: implicit versus explicit representation, offline versus online updates, and short-term versus long-term persistence. It also formalizes memory writing, routing, state transitions, and consolidation, while examining hybrid architectures, system-level efficiency trade-offs, and multidimensional evaluation. The paper aims to connect transient attention, recurrent state dynamics, parameter-efficient adaptation, and scalable lookup storage within a common framework for designing scalable and adaptive language models.
No heat snapshots are available in the last 24 hours.