Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Can Foundation Models Moderate Content? Instruction- vs. Example-Driven Policy Evaluation

First seen · 9/10/2026, 12:30 AMLatest activity · 9/10/2026, 12:30 AM

Evaluating content moderation through Vision-Language Models, this study systematically benchmarks instruction-driven policy interpretation against example-driven precedent matching. Using ModerationBench—a dataset of 4,000 manually annotated, in-the-wild posts from the Bluesky platform—the authors demonstrate that modern foundation models nearly triple the F1 performance of the network's deployed moderation stack (0.60 versus 0.22). Both paradigms achieve comparable peak efficacy, indicating that multimodal foundation models can reliably ground fluid community standards in either formal rules or historical precedents.

Event heat · last 24 hours

There are 6 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 20:00; latest heat is 0.

There are 6 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 20:00; latest heat is 0.10.509/12, 20:00, event heat 09/12, 23:00, event heat 09/13, 02:00, event heat 09/13, 05:00, event heat 09/13, 08:00, event heat 09/13, 11:00, event heat 024 hours agoNow
  1. 9/12, 20:00, event heat 0
  2. 9/12, 23:00, event heat 0
  3. 9/13, 02:00, event heat 0
  4. 9/13, 05:00, event heat 0
  5. 9/13, 08:00, event heat 0
  6. 9/13, 11:00, event heat 0

Reporting Timeline

  1. AggregatorarXiv9/10, 12:30 AMnot independentRepresentative
    Can Foundation Models Moderate Content? Instruction- vs. Example-Driven Policy Evaluation