The paper introduces Vera, an end-to-end automated safety-testing framework for nondeterministic LLM agents. Its pipeline discovers emerging risks from the literature, composes executable safety cases across risk, attack, and environment taxonomies, and adaptively executes agents in isolated sandboxes. Verification relies on observable environment state and tool-call evidence rather than agent self-report. Evaluated on OpenClaw, Hermes, Codex, and Claude Code, Vera reports average attack success rates of up to 93.9% under multi-channel attacks. The authors also release Vera-Bench, containing 1,600 executable cases across 124 risk categories and three execution settings.
No heat snapshots are available in the last 24 hours.