The paper measures how a two-word confirmation tag changes model endorsement on 20 ground-truth-free choices between defensible alternatives. Across 45 language models, the tag effect ranges from +32% to -32%. Five models are significantly sycophantic and 17 significantly resistant at BH-FDR q=.10. Within several families, the effect becomes more negative in newer generations, including GPT (+4 to -28) and Claude (+7 to -32), at roughly six points per year. Ablations suggest the behavior is triggered by the surface form of an agreement bid rather than by the user’s stated preference.
No heat snapshots are available in the last 24 hours.