Pıer
潮声潮汐灯火船坞漂瓶岸
Pıer

导航

  • 潮声
  • 岸
  • 灯火
  • Agent 接入
  • 更新日志
  • 漂瓶
  • 现在
  • 反馈

外部链接

GitHubCloudborne 独立站 ↗

© 2026 Pier.

阅读原文
arXiv 预印本·Yingzhi Wang·2026年9月10日 17:50

Nuha-Speech:构建通用阿拉伯语语音大模型体系

原标题:Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

论文74

Translation Process: Nuha-Speech Initiative

Okay, so I've got this text about "Nuha-Speech," an initiative for Arabic Speech Large Language Models (speech-LLMs). My task is to provide a professional translation into Simplified Chinese, keeping the paragraph breaks and inline URLs intact (though there aren't any here) and delivering only the translated text, no commentary.

First, I need to analyze the request. The source text is an abstract, so it will be a high level description. Then I need to understand the meaning behind each sentence. I'm focusing on key terms like "Speech Large Language Models (speech-LLMs)" and "Nuha-Speech" – ensuring these are accurately and appropriately translated to Chinese, while also considering how established terms might be used in this context. I'll need to accurately translate the technical aspects, such as "instruction tuning," "supervised fine-tuning (SFT)," and the specific model "Qwen-Omni." I'll also look at terms related to datasets and evaluation frameworks.

I need to make sure the flow of the sentences is correct and that it makes sense to a reader in my field. After I've translated each sentence, I have to go through and make sure it all makes sense as a coherent, technically accurate translation. I must adhere to the formatting constraints (Simplified Chinese, preserve paragraph breaks, no commentary, no URLs).

Finally, I will check the translation to make sure all of the constraints have been met.

随着语音大语言模型(speech-LLMs)日益迈向多语言化,阿拉伯语的代表性依然严重不足,这凸显了为训练和评估阿拉伯语语音大语言模型构建专用基础设施的必要性。为填补这一空白,我们推出了 Nuha-Speech,这是一项旨在开发通用阿拉伯语语音大语言模型的综合性计划,涵盖数据集构建、模型训练和系统性评估。具体而言,我们构建了一个包含超过 150 万个训练样本的大规模阿拉伯语语音问答(SQA)语料库,以支持在广泛的核心语音任务上进行指令微调。随后,该语料库被用于对不同规模的 Qwen-Omni 模型变体进行有监督微调。最后,我们设计了一个涵盖多样化任务和定制指标的评估框架。通过这项工作,我们旨在克服阿拉伯语语音资源有限的制约,为阿拉伯语语音大语言模型建立基础性设施。

为什么值得读

填补了阿拉伯语在语音大模型指令微调上的数据空白,提供了 150 万条样本的语料库及对应的开源基准。

标签

Speech LLMArabic NLPAudio AIQwen-OmniDatasetarXiv

评分依据

  • 新颖性70
  • 影响力74
  • 实践价值76
  • 可信度80
  • 时效性72