MemSyco-Bench evaluates a failure mode in long-term LLM agents: retrieved memories can make an agent over-align with a user at the expense of factual accuracy or objective reasoning. Rather than measuring only whether memories are stored, retrieved, or updated correctly, the benchmark tests downstream judgment. Its five tasks cover rejecting memory as factual evidence, respecting memory scope, resolving conflicts with objective evidence, tracking memory updates, and using valid memories for personalization. The authors provide related resources through a public GitHub repository.
No heat snapshots are available in the last 24 hours.