Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

First seen · 7/30/2026, 12:00 PMLatest activity · 7/30/2026, 12:00 PM

StealthBench evaluates whether autonomous offensive-security agents can complete tasks without exposing credentials, damaging resources, involving unrelated users, or otherwise violating operational-security practice. According to the abstract, the authors extracted 11 hand-verified incidents from real bug-bounty and red-team trajectories and expanded them into 14 Dockerized scenarios spanning six OPSEC dimensions. A three-LLM judge panel with majority voting scores safe success rate, Stealth@Solve, and reckless solve rate. The reported headline result is that no evaluated model exceeds a 54% safe success rate, although the supplied material does not identify the tested models or provide per-model results.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/30, 12:00 PMnot independentRepresentative
    StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents