This paper presents a systematic empirical study of sycophancy in LLM-based code smell detection. Using the MLCQ dataset, it tests confirmation-biased framings, contradictory hints, and false premises. The reported Decision Flip Rate reaches 72%, while False Alignment Rate exceeds 90%, suggesting that model outputs can be driven by prompt assumptions rather than code evidence. The proposed Evidence-Guided Debiasing Prompting (EGDP) strategy enforces evidence-first reasoning and reportedly reduces these rates to as low as 12% and 21%, respectively, while increasing reliance on structural evidence.
No heat snapshots are available in the last 24 hours.