GigaChat Audio is a time-aware audio language model designed to answer questions over recordings of up to 120 minutes while producing explicit timestamps. The method interleaves periodic time markers with continuous audio tokens and uses large-scale synthetic supervision generated by a cascaded pipeline. The paper reports temporal-grounding results on short- and long-audio benchmarks, supports timestamped fragment descriptions and summaries, and studies the effects of time representation, marker frequency, tokenization, and duration mixtures through ablations. Model weights and datasets are released on Hugging Face as GigaChat3.1-Audio-10B-A1.8B.
No heat snapshots are available in the last 24 hours.