The paper introduces BioSecBench-Refusal, a benchmark that evaluates both task performance and refusal behavior for AI agents operating on biological research workflows. It pairs 61 legitimate Routine tasks adapted from published literature with 46 fictional Red-Team tasks that resemble real research while concealing a biosecurity hazard. Across 16 model-harness configurations, refusal rates were 7% to 74% on Routine tasks and 1% to 62% on Red-Team tasks. The authors report that provider API filters often triggered refusals before agentic reasoning, while models allowed to reason had greater potential to detect genuine threats.
No heat snapshots are available in the last 24 hours.