Audio-Zero presents a label-free self-evolution framework for fine-grained reasoning in large audio-language models. It builds an auditory self-play game from unlabeled contrastive audio pairs: most players hear a reference clip while one odd listener hears a subtle variant. The model generates auditory clues, then identifies the odd listener by reasoning over inconsistencies among those clues. Because the odd listener is known by construction, the game supplies verifiable rewards without annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini, and MMAR reportedly improve fine-grained reasoning while preserving broad audio understanding.
No heat snapshots are available in the last 24 hours.