This paper evaluates security agents using both task success and operational cost across offensive Cybench challenges and defensive Splunk BOTS v1 investigations. The authors report different scaling regimes: offensive CTF performance generally improves with additional test-time compute, allowing scaled open-weight models to approach frontier proprietary systems at competitive cost. Defensive SOC investigation shows weaker scaling with raw reasoning budget and depends more on disciplined tool use, telemetry navigation, and selective enrichment. The paper argues that security benchmarks should report economic efficiency and operational fit, not only peak capability, and provides an interactive results site.
No heat snapshots are available in the last 24 hours.