This paper argues that the quality–intelligibility trade-off in generative streaming Target Speaker Extraction (TSE) is caused primarily by a poor optimization anchor rather than by streaming constraints. It enlarges the Conformer convolution kernel for richer local spectro-temporal modeling and uses WavLM cosine similarity to rank preference pairs for Direct Preference Optimization (DPO). With a 560 ms streaming chunk size, the reported word error rate improves from 0.138 to 0.123, a 10.9% relative intelligibility gain, while audio quality and speaker similarity also improve marginally.
No heat snapshots are available in the last 24 hours.