Article robust multimodal emotion recognition from conversation with transformer-based crossmodality the title fusion

104Citations
Citations of this article
84Readers
Mendeley users who have this article in their library.

Abstract

Decades of scientific research have been conducted on developing and evaluating methods for automated emotion recognition. With exponentially growing technology, there is a wide range of emerging applications that require emotional state recognition of the user. This paper investigates a robust approach for multimodal emotion recognition during a conversation. Three separate models for audio, video and text modalities are structured and fine-tuned on the MELD. In this paper, a transformer-based crossmodality fusion with the EmbraceNet architecture is employed to estimate the emotion. The proposed multimodal network architecture can achieve up to 65% accuracy, which significantly surpasses any of the unimodal models. We provide multiple evaluation techniques applied to our work to show that our model is robust and can even outperform the state-of-the-art models on the MELD.

Cite

CITATION STYLE

APA

Xie, B., Sidulova, M., & Park, C. H. (2021). Article robust multimodal emotion recognition from conversation with transformer-based crossmodality the title fusion. Sensors, 21(14). https://doi.org/10.3390/s21144913

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free