The paper reframes safety checking for LLM tool calls from binary safe/unsafe classification to per-instance routing among EXECUTE, ASK, and REFUSE. It introduces Safety Sentry, a lightweight guard model whose inference requires a single decoding call. One decoding-time threshold can reposition the same checkpoint for deployments with different risk tolerances without retraining. The abstract reports higher overall accuracy and safety-related recall than a broad set of open-weight and frontier closed-source baselines, while controlling both directional error rates. However, the supplied record contains no benchmark names, numerical results, model size, or experimental details.
No heat snapshots are available in the last 24 hours.