Trimming audio-encoder depth cuts inference overhead in speech LLMs, but directly dropping blocks often perturbs decoder embeddings, triggering token deletions and premature end-of-sequence stops. X-AuT addresses this with a progressive compression framework that selects layer subsets through behavioral probes and recovers performance via cross-scale distillation and LoRA tuning while keeping the language model backbone frozen. Evaluated on Qwen3-ASR-0.6B across ten Chinese-English benchmarks, pruning the audio encoder from 18 to 14 layers reduced audio-tower parameters by 20.7% with a macro-average error of 5.75%, outpacing direct pruning.
There are 6 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 20:00; latest heat is 0.