RetroThinker: Enabling Retrospective Thinking in Speech LLMs
First seen · 9/11/2026, 01:41 AMLatest activity · 9/11/2026, 01:41 AM
Streaming SpeechLLMs face a severe accuracy-latency dilemma when tackling complex reasoning under conversational time constraints. RetroThinker introduces a multi-stage post-training framework that allows Moshi to self-verify and forward-correct reasoning steps on the fly while listening to user input. By combining retrospective supervised fine-tuning with length-based direct preference optimization, the system achieves an 11% absolute accuracy gain on the GSM8K benchmark without worsening interaction latency.
Event heat · last 24 hours
There are 7 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 11:00; latest heat is 0.