Evaluating content moderation through Vision-Language Models, this study systematically benchmarks instruction-driven policy interpretation against example-driven precedent matching. Using ModerationBench—a dataset of 4,000 manually annotated, in-the-wild posts from the Bluesky platform—the authors demonstrate that modern foundation models nearly triple the F1 performance of the network's deployed moderation stack (0.60 versus 0.22). Both paradigms achieve comparable peak efficacy, indicating that multimodal foundation models can reliably ground fluid community standards in either formal rules or historical precedents.
There are 6 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 20:00; latest heat is 0.