This paper introduces self-state attacks, in which a self-hosted AI agent is compromised by modifying its own memory or configuration through legitimate operating-system system calls. The authors define a four-axis attack space covering target, mechanism, granularity, and temporal behavior. They collect live traces from a representative agent under distinct workload profiles, instantiate 23 attack cells, and inject 43 concrete operations into those traces. Their evaluation finds that layered defenses can cover most cells: access control for instruction and configuration files, workload-conditioned detection for memory, and periodic backups for recovery. However, some attacks remain structurally indistinguishable at the OS level.
No heat snapshots are available in the last 24 hours.