Nemotron Labs introduces Nemotron-Labs-Audex-30B-A3B, a unified audio-text mixture-of-experts model built on Nemotron-Cascade-2-30B-A3B. A single Transformer decoder handles projected audio inputs, text tokens, and quantized audio output tokens in one generation space. Training uses 157.4B audio tokens and 320.5B text tokens, followed by supervised training, text-only Cascade RL, and multi-domain on-policy distillation. The abstract reports leading results across audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, with marginal or no regression in text reasoning and agentic capabilities. Checkpoints are released for research.
No heat snapshots are available in the last 24 hours.