This paper proposes an evaluation protocol for AI penetration-testing agents that prioritizes validated vulnerability discovery over predefined task completion. The protocol targets complex systems with multiple attack surfaces and vulnerability classes, combining structured ground truth, LLM-based semantic matching, bipartite resolution for ambiguous findings, continuous ground-truth maintenance, repeated and cumulative evaluation of stochastic agents, efficiency metrics, and reduced-suite selection. The authors also release expert-annotated ground truth and implementation code through the EthiBench repository.
No heat snapshots are available in the last 24 hours.