Abstract
People generally perceive other people’s emotions based on speech and facial expressions, so it can be helpful to use speech signals and facial images simultaneously. However, because the characteristics of speech and image data are different, combining the two inputs is still a challenging issue in the area of emotion-recognition research. In this paper, we propose a method to recognize emotions by synchronizing speech signals and image sequences. We design three deep networks. One of the networks is trained using image sequences, which focus on facial expression changes. Facial landmarks are also input to another network to reflect facial motion. The speech signals are first converted to acoustic features, which are used for the input of the other network, synchronizing the image sequence. These three networks are combined using a novel integration method to boost the performance of emotion recognition. A test comparing accuracy is conducted to verify the proposed method. The results demonstrated that the proposed method exhibits more accurate performance than previous studies.
Author supplied keywords
Cite
CITATION STYLE
Byun, S. W., & Lee, S. P. (2021). Human emotion recognition based on the weighted integration method using image sequences and acoustic features. Multimedia Tools and Applications, 80(28–29), 35871–35885. https://doi.org/10.1007/s11042-020-09842-1
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.