A Survey of Multi-modal Emotion Recognition Based on Deep Learning

  • Jia M
  • Sun Z
N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Multi-modal emotion recognition technology explores emotion recognition by integrating facial expression, voice intonation, text analysis and other multi-source data, so as to improve the naturalness and accuracy of human-computer interaction. Aiming at the emerging field of multi-modal emotion recognition, this paper introduces three single-modal emotion recognition methods of text, face and voice, especially the problem of multi-modal emotion fusion, and introduces the methods with high success rate of multi-modal emotion fusion recognition in recent years. Through comparative analysis, the conclusion is drawn that the current fusion methods are more complicated and the fusion success rate has been improved to some extent. However, the number of data sets on multi-modal emotion analysis is small, and the research on gesture and other modes of emotion recognition is also scarce. In the later stage, it is necessary to enrich the data set and add new modes to improve the accuracy and robustness of the multi-modal emotion recognition and analysis system.

Cite

CITATION STYLE

APA

Jia, M., & Sun, Z. (2024). A Survey of Multi-modal Emotion Recognition Based on Deep Learning. Highlights in Science, Engineering and Technology, 119, 533–540. https://doi.org/10.54097/37zncv36

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free