Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

First seen · 8/1/2026, 08:00 PMLatest activity · 8/1/2026, 08:00 PM

This paper studies internal representations of agentic LLMs exposed to indirect prompt injection (IPI), such as malicious side tasks embedded in tool outputs. Across six models, linear probes over pre-generation hidden states reportedly detect IPI exposure with over 90% AUROC on unseen attacks, instructions, and task suites, including cross-lingual and adaptive settings. The authors introduce AGRI, a probe-gated reasoning defense that activates anti-injection reasoning when needed. On difficult AgentDojo evaluations, AGRI reduces Qwen3.5-27B's attack success rate from 34.6% to 0% while largely preserving clean-task utility. They also analyze model-specific natural-language explanations associated with the latent signals.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv8/1, 08:00 PMnot independentRepresentative
    Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure