Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

First seen · 7/16/2026, 12:00 PMLatest activity · 7/16/2026, 12:00 PM

This paper studies image safety guardrails that must follow a supplied policy rather than treat safety as an intrinsic property of an image. It introduces PolicyShiftBench, containing 2,000 policy-discriminative instances over 265 images, with 7.55 policy-conditioned prompts per image on average. The authors propose the 7B PolicyShiftGuard, trained with Randomized Policy SFT (RP-SFT) and Boundary-Pair Policy Adaptation (BP-Adapt). According to the abstract, it reaches 76.9 Avg. F1 and 72.1 Avg. PSS, transfers to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off through concise outputs.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/7, 03:04 PMnot independent
    PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
  2. AggregatorHuggingFace Daily Papers7/16, 12:00 PMnot independentRepresentative
    PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails