This paper studies internal representations of agentic LLMs exposed to indirect prompt injection (IPI), such as malicious side tasks embedded in tool outputs. Across six models, linear probes over pre-generation hidden states reportedly detect IPI exposure with over 90% AUROC on unseen attacks, instructions, and task suites, including cross-lingual and adaptive settings. The authors introduce AGRI, a probe-gated reasoning defense that activates anti-injection reasoning when needed. On difficult AgentDojo evaluations, AGRI reduces Qwen3.5-27B's attack success rate from 34.6% to 0% while largely preserving clean-task utility. They also analyze model-specific natural-language explanations associated with the latent signals.
No heat snapshots are available in the last 24 hours.