Abstract
Speech Emotion Recognition (SER) plays an important role in the field of Human-Computer Interaction (HCI). With the rapid development of machine learning technology, SER based on deep learning has also made new progress. At the same time, the insufficient performance of single feature leads to a lower recognition accuracy. Aiming at this problem, in this study, a multi-feature fusion method is proposed to improve the accuracy of SER. Firstly, several different spectrogram features are divided into three sub-spectrograms along frequency to make the local features more prominent. Secondly, each independent sub-spectrogram is used as input of the improved Residual Network (ResNet) structure based on attention mechanism to extract emotion feature maps. Finally, the emotion feature maps are fused in the feature layer which is connected to the output layer. The experimental results show that the mentioned method can better classify kinds of emotions. Specifically, an accuracy of 98% is achieved on the ESD dataset, which is nearly 6 percentage points higher than other methods. In order to verify the effectiveness of the method, we also carried out experiments on the SER standard dataset IEMOCAP, and the accuracy rate reached 76.54%. Proving the effectiveness of the mentioned method.
Author supplied keywords
Cite
CITATION STYLE
Guo, Y., Zhou, Y., Xiong, X., Jiang, X., Tian, H., & Zhang, Q. (2023). A Multi-Feature Fusion Speech Emotion Recognition Method Based on Frequency Band Division and Improved Residual Network. IEEE Access, 11, 86013–86024. https://doi.org/10.1109/ACCESS.2023.3299822
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.