StealthBench evaluates whether autonomous offensive-security agents can complete tasks without exposing credentials, damaging resources, involving unrelated users, or otherwise violating operational-security practice. According to the abstract, the authors extracted 11 hand-verified incidents from real bug-bounty and red-team trajectories and expanded them into 14 Dockerized scenarios spanning six OPSEC dimensions. A three-LLM judge panel with majority voting scores safe success rate, Stealth@Solve, and reckless solve rate. The reported headline result is that no evaluated model exceeds a 54% safe success rate, although the supplied material does not identify the tested models or provide per-model results.
No heat snapshots are available in the last 24 hours.