Voice Memory is an inference-only scheme in which a frozen speech corrector reads one auditable memory.md per domain and decides, utterance by utterance, whether to correct an ASR hypothesis or retain the 1-best output. An asynchronous optimizer makes bounded memory edits and accepts only those that strictly improve a held-out score. The authors report that, across ten HyPoradise domains, weighted word error rate falls from 8.36% to 7.52%, or 7.47% with three additional in-context examples, without any dataset dropping below its 1-best baseline.
No heat snapshots are available in the last 24 hours.