The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
This paper studies over-inference (OI), where personalized LLMs fabricate user attributes beyond the available evidence. The authors introduce MirageBench with 150 personas, six personalization tasks, a four-way faithfulness taxonomy, and 143,616 judged claims across 12 models from seven families. Every evaluated model over-inferred 35%–49% of claims, with a cross-model mean of 41.6%. Model self-assessments were negatively rank-correlated with judge-measured OI, suggesting that self-reported confidence is unreliable for comparing personalization systems.
Why it's worth reading
Persistent memory is moving into mainstream products, while this evaluation shows that models can systematically invent user traits and that their own self-assessments may mislead model selection.