Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

Read original
arXiv·Yingzhi Wang·Sep 10, 2026, 5:50 PM

Nuha-Speech: Establishing Foundations for Arabic Speech-LLMs

Original title:Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Papers74

While multimodal speech-LLMs advance rapidly across major languages, Arabic remains constrained by a persistent shortage of speech-text training assets. The Nuha-Speech project establishes an end-to-end infrastructure to bridge this gap, introducing an Arabic Speech Question-Answering corpus with over 1.5 million instruction samples. Leveraging Qwen-Omni model variants across scales alongside a dedicated evaluation benchmark, the work provides a grounded reference for developing conversational speech models in resource-limited languages.

Why it's worth reading

It addresses the data shortage in Arabic speech AI with a 1.5M-sample QA corpus and standardized fine-tuning baselines on Qwen-Omni.

Tags

Speech LLMArabic NLPAudio AIQwen-OmniDatasetarXiv

Score breakdown

  • Novelty70
  • Impact74
  • Practicality76
  • Credibility80
  • Timeliness72