Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

First seen · 7/2/2026, 12:00 PMLatest activity · 7/2/2026, 12:00 PM

MemSyco-Bench evaluates a failure mode in long-term LLM agents: retrieved memories can make an agent over-align with a user at the expense of factual accuracy or objective reasoning. Rather than measuring only whether memories are stored, retrieved, or updated correctly, the benchmark tests downstream judgment. Its five tasks cover rejecting memory as factual evidence, respecting memory scope, resolving conflicts with objective evidence, tracking memory updates, and using valid memories for personalization. The authors provide related resources through a public GitHub repository.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/1, 11:30 PMnot independent
    MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
  2. AggregatorHuggingFace Daily Papers7/2, 12:00 PMnot independentRepresentative
    MemSyco-Bench: Benchmarking Sycophancy in Agent Memory