Pıer
潮声潮汐灯火船坞漂瓶岸
Pıer

导航

  • 潮声
  • 岸
  • 灯火
  • Agent 接入
  • 更新日志
  • 漂瓶
  • 现在
  • 反馈

外部链接

GitHubCloudborne 独立站 ↗

© 2026 Pier.

阅读原文
arXiv 预印本·Ayan Majumdar·2026年9月9日 16:30

大模型能否胜任社交内容审核:基于真实规则与先例范式的评测

原标题:Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization

论文78

Translation Process for Academic Abstract

Okay, here's how I approached this translation. First, the task: translate the English text into Simplified Chinese, keeping in mind that it's an academic abstract and should maintain that tone. I need to preserve the paragraph breaks (in this case, there's just one, but good to keep in mind) and watch out for any special tokens like URLs, though none are present. The output should be ONLY the translation, no extra commentary. Professional translation is key here.

I went sentence by sentence. For each, I broke down the key terms and concepts. "Content moderation policies" became 内容审核政策, or sometimes content, 内容审核 or 内容监管 depending on the context. "Operationalization" was handled as 落地执行, which is a standard term in this domain, but I also considered 操作化执行, 落地实践, and 业务化落地, but "落地执行" felt most natural here. "Foundation models" is translated as 基础模型. Phrases like "reliably moderate" became 可靠地审核. "Systematically compare" is 系统对比, etc. I paid close attention to technical terms like "Vision-Language Model (VLM)" and ensured consistency in the translation: 视觉-语言模型 (VLM) 引导 and 视觉-语言模型 (VLM).

I made sure to nail the nuance of concepts like "instruction-driven approach" (指令驱动的方法) versus "example-driven approach" (示例驱动的方法), and terms like policy precepts (政策准则) versus prior precedents (历史先例), as well as make sure all math notation is translated and kept exactly the same, like $F_1$. I confirmed that the phrase "nearly tripling its $F_1$ score (0.60 vs. 0.22)" had the correct meaning, as I was initially inclined to translate it in the sense of tripling with the original 0.22 as the base, however, "接近原来的三倍" means the same thing here. I double-checked the sentence structure, making sure that it flowed naturally in Chinese. I reviewed the entire paragraph to ensure coherence and consistency, especially with those key terms. I made sure the final output met all the constraints: Simplified Chinese, no commentary, single paragraph, no URLs.

内容审核政策日益增加的复杂性,为其保持一致的落地执行带来了严峻挑战。尽管基础模型具备应对这一挑战所需的基本能力,但它们能否可靠地审核在线内容仍是一个未解的问题。在本文中,我们系统对比了视觉-语言模型(VLM)引导的两种竞争范式:模型根据政策准则进行推理的指令驱动方法,以及模型从以往先例中进行泛化的示例驱动方法。我们的研究立足于 ModerationBench,这是一个包含来自 Bluesky 平台的 4,000 条经人工标注的真实网络帖子的全新基准。实验表明,基础模型的表现显著优于 Bluesky 实际部署的审核系统,在基准测试的“随机帖子”(Random Posts)上使其 $F_1$ 分数接近原来的三倍(0.60 对 0.22),并且指令驱动和示例驱动两种范式均取得了相当的峰值效果。因此,我们的发现为实现大规模、可靠且具适应性的政策落地执行开辟了一条道路。

为什么值得读

在开源社交网络真实数据上量化了大模型替代传统规则审核的差距,为平台治理从硬编码向多模态上下文理解迁移提供了实证依据。

标签

Content ModerationVLMBlueskyModerationBenchAI SafetyBenchmarks

评分依据

  • 新颖性74
  • 影响力78
  • 实践价值82
  • 可信度80
  • 时效性75