This paper presents a lightweight and explainable framework for speech emotion recognition. It combines log-Mel spectrogram inputs, a compact convolutional neural network, and attentive statistics pooling to emphasize emotionally salient temporal regions. Grad-CAM is used to visualize the time-frequency areas influencing predictions. On the SAVEE emotional speech dataset, the authors report competitive recognition performance with substantially fewer parameters than many existing SER architectures. The work targets a practical balance among accuracy, computational efficiency, and interpretability for applications such as healthcare, customer service, and human-computer interaction.
No heat snapshots are available in the last 24 hours.