Necessary or Sufficient? Evaluating LLM Explanations with Behavioural Evidence
Original title:Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
When LLMs make decisions within agent workflows, they often cite the primary factors behind their judgements. Evaluating eight models across the Claude, GPT, and Gemini families through controlled black-box interventions, this study tests whether these cited explanations meet behavioural standards of necessity and sufficiency. In an advisor-recommendation benchmark, uncited factors held more measurable sway than the lowest-ranked cited factor in over 57% of cases. The findings indicate that while self-reported explanations carry partial signal, they cannot be reliably trusted as true causal accounts for safety auditing.
Why it's worth reading
As agent systems increasingly rely on self-generated explanations for oversight and error diagnosis, this study provides timely empirical evidence that self-reported factors often fail to capture the true causal drivers of LLM decisions.