This paper argues that verbatim and intended transcription styles become an uncontrolled latent variable when ASR models learn from heterogeneous annotations. The authors report that style mismatch can account for up to 60% of reported WER. Coverage-aware decoder task tokens trained on parallel verbatim/intended transcripts raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. English-only fine-tuning improves verbatim accuracy, disfluency detection, and intended-mode quality across English and German. Supervised cross-attention fine-tuning also improves word-level timestamps on disfluent speech, and the authors introduce verbatimize for scalable corpus enrichment.
No heat snapshots are available in the last 24 hours.