Abstract
In order to establish the mutual fusion relationship between the global and timing features of speech and achieve better speech emotion recognition, this paper proposes a PAD 3-D space emotion recognition method based on feature fusion. This paper uses the Opensmile to extract four speech global feature sets: IS09-emotion, IS10-paraling, IS11-speaker , IS12-speaker-trait, and establishes a deep learning model combined with CNN, LSTM, and attention mechanism to extract the timing features of each frame simultaneously. Finally, the global features and timing features are fused through the Stacking model. The experimental results show that compared with using global features or timing features alone for speech emotion recognition, the method of using the Stacking model to fuse global features and temporal features can effectively improve the accuracy of speech emotion recognition.
Cite
CITATION STYLE
Pengfei, X., Houpan, Z., & Weidong, Z. (2020). PAD 3-D speech emotion recognition based on feature fusion. In Journal of Physics: Conference Series (Vol. 1616). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/1616/1/012106
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.