The paper introduces GigaAM Multilingual, a Conformer-based multilingual ASR foundation model focused on underrepresented Central Asian languages, including Kazakh, Kyrgyz, and Uzbek. It is pretrained on 2 million hours of audio with a HuBERT-style objective, using cluster-level balancing to reduce head-language dominance. During fine-tuning, domain-aware sampling is used to improve adaptation under data imbalance. The supplied abstract reports gains over Whisper Large v3 and Omnilingual-1B on the target languages, particularly spontaneous speech, while retaining efficiency. The encoder and ASR model are released.
No heat snapshots are available in the last 24 hours.