Shieldstral is a 3B-parameter, policy-adaptive multimodal safety classifier that casts content moderation as binary question answering. This formulation allows datasets with different safety taxonomies to be consolidated into one training framework. The paper describes the curation and generation of roughly 54.1 million samples and introduces a fine-grained evaluation set for policy adaptability. According to the supplied abstract, Shieldstral matches or exceeds models nearly seven times larger on text safety benchmarks and establishes a new state of the art in multimodal safety classification.
No heat snapshots are available in the last 24 hours.