This paper presents a multimodal framework for binary sentiment polarity classification from speech. It generates transcripts with automatic speech recognition, translates them into multiple languages, and progressively fuses audio and multilingual text through cascaded cross-modal Transformer blocks. Knowledge from this multimodal teacher is distilled into an audio-only student model. The authors report that both automatically generated transcripts and translations improve performance, while distillation enhances the audio-only model without adding inference-time computational cost. The abstract does not specify the dataset, metrics, or numerical gains.
No heat snapshots are available in the last 24 hours.