This paper studies image safety guardrails that must follow a supplied policy rather than treat safety as an intrinsic property of an image. It introduces PolicyShiftBench, containing 2,000 policy-discriminative instances over 265 images, with 7.55 policy-conditioned prompts per image on average. The authors propose the 7B PolicyShiftGuard, trained with Randomized Policy SFT (RP-SFT) and Boundary-Pair Policy Adaptation (BP-Adapt). According to the abstract, it reaches 76.9 Avg. F1 and 72.1 Avg. PSS, transfers to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off through concise outputs.
No heat snapshots are available in the last 24 hours.