arXivYi-Jen Shih
RetroThinker: Enabling Retrospective Thinking in Speech LLMs
Papers81
Streaming SpeechLLMs face a severe accuracy-latency dilemma when tackling complex reasoning under conversational time constraints. RetroThinker introduces a multi-stage post-training framework that allows Moshi to self-verify and forward-correct reasoning steps on the fly while listening to user input. By combining retrospective supervised fine-tuning with length-based direct preference optimization, the system achieves an 11% absolute accuracy gain on the GSM8K benchmark without worsening interaction latency.
Why it's worth reading
As speech models evolve toward complex reasoning, this paper offers a practical post-training framework for dynamic, on-the-fly error correction without introducing conversational lag.
Tags
SpeechLLMReasoningChain-of-ThoughtMoshiDPOStreaming AI