LazyMem postpones memory construction until query time. It first retrieves a broad candidate pool, then uses a lightweight model over overlapping parallel windows to select and compress only content relevant to the current query. Trained with supervised fine-tuning and reinforcement learning, LazyMem-4B reaches 0.85 LLM-judge accuracy on LongMemEval while using 213 answer-context memory tokens, 21.0 times fewer than the strongest non-oracle baseline. The method also generalizes to LoCoMo without target-domain training and lowers mean latency compared with a prior query-time baseline.
No heat snapshots are available in the last 24 hours.