This paper introduces HIVE, the Human Input-Variation Engine, for testing instruction-tuned models under voice-transcription and QWERTY-keyboard perturbations. The authors report that voice transcription consistently reduces accuracy more than keyboard noise, with structural rewriting rather than filler words causing most of the damage. Their proposed common explanation is token survival: corrupting original question tokens hurts, while adding words alongside preserved tokens costs relatively little. The reported channel gap appears on constructed or deduced answers, not multiple choice. Lightweight adaptation does not remove it, while additional thinking budget largely repairs keyboard noise but not spoken input.
No heat snapshots are available in the last 24 hours.