Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingNewsWatching0 independent reports0

GigaAM Multilingual: A Foundation Model for Underrepresented Languages

First seen · 7/21/2026, 12:00 PMLatest activity · 7/21/2026, 12:00 PM

The paper introduces GigaAM Multilingual, a Conformer-based multilingual ASR foundation model focused on underrepresented Central Asian languages, including Kazakh, Kyrgyz, and Uzbek. It is pretrained on 2 million hours of audio with a HuBERT-style objective, using cluster-level balancing to reduce head-language dominance. During fine-tuning, domain-aware sampling is used to improve adaptation under data imbalance. The supplied abstract reports gains over Whisper Large v3 and Omnilingual-1B on the target languages, particularly spontaneous speech, while retaining efficiency. The encoder and ASR model are released.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/21, 12:00 PMnot independentRepresentative
    GigaAM Multilingual: A Foundation Model for Underrepresented Languages