Can Foundation Models Moderate Content? Instruction- vs. Example-Driven Policy Evaluation
Original title:Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization
Evaluating content moderation through Vision-Language Models, this study systematically benchmarks instruction-driven policy interpretation against example-driven precedent matching. Using ModerationBench—a dataset of 4,000 manually annotated, in-the-wild posts from the Bluesky platform—the authors demonstrate that modern foundation models nearly triple the F1 performance of the network's deployed moderation stack (0.60 versus 0.22). Both paradigms achieve comparable peak efficacy, indicating that multimodal foundation models can reliably ground fluid community standards in either formal rules or historical precedents.
Why it's worth reading
It provides rigorous benchmark evidence on in-the-wild social data, showing how multimodal foundation models substantially outperform current deployed platform moderation using either rule precepts or precedent examples.