Read original
HuggingFace Daily PapersYushi SunPapers88

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

This paper studies over-inference (OI), where personalized LLMs fabricate user attributes beyond the available evidence. The authors introduce MirageBench with 150 personas, six personalization tasks, a four-way faithfulness taxonomy, and 143,616 judged claims across 12 models from seven families. Every evaluated model over-inferred 35%–49% of claims, with a cross-model mean of 41.6%. Model self-assessments were negatively rank-correlated with judge-measured OI, suggesting that self-reported confidence is unreliable for comparing personalization systems.

Why it's worth reading

Persistent memory is moving into mainstream products, while this evaluation shows that models can systematically invent user traits and that their own self-assessments may mislead model selection.

Tags

个性化LLM用户画像模型记忆事实性评测基准AI安全MirageBench