When LLMs make decisions within agent workflows, they often cite the primary factors behind their judgements. Evaluating eight models across the Claude, GPT, and Gemini families through controlled black-box interventions, this study tests whether these cited explanations meet behavioural standards of necessity and sufficiency. In an advisor-recommendation benchmark, uncited factors held more measurable sway than the lowest-ranked cited factor in over 57% of cases. The findings indicate that while self-reported explanations carry partial signal, they cannot be reliably trusted as true causal accounts for safety auditing.
No heat snapshots are available in the last 24 hours.