PAD 3-D speech emotion recognition based on feature fusion

5Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

In order to establish the mutual fusion relationship between the global and timing features of speech and achieve better speech emotion recognition, this paper proposes a PAD 3-D space emotion recognition method based on feature fusion. This paper uses the Opensmile to extract four speech global feature sets: IS09-emotion, IS10-paraling, IS11-speaker , IS12-speaker-trait, and establishes a deep learning model combined with CNN, LSTM, and attention mechanism to extract the timing features of each frame simultaneously. Finally, the global features and timing features are fused through the Stacking model. The experimental results show that compared with using global features or timing features alone for speech emotion recognition, the method of using the Stacking model to fuse global features and temporal features can effectively improve the accuracy of speech emotion recognition.

Cite

CITATION STYLE

APA

Pengfei, X., Houpan, Z., & Weidong, Z. (2020). PAD 3-D speech emotion recognition based on feature fusion. In Journal of Physics: Conference Series (Vol. 1616). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/1616/1/012106

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free