A Complementary Fusion Framework for Robust Multimodal Emotion Recognition

0Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

This paper presents a novel dual-stream framework for multimodal emotion recognition, engineered to address the varying complexity inherent in emotional expressions. The proposed architecture uniquely integrates a graph embedding as an auxiliary modality to explicitly model temporal correlations between utterances, and processes features through two complementary sub-models operating in parallel: a Cross-Attention Transformer Mixture-of-Experts (MoE) model and a Sum-Product Linear MoE model. The former deciphers nuanced and ambiguous emotions, such as ‘fear’ and ‘sadness’, by leveraging a deep cross-attention mechanism to model intricate, bidirectional dependencies between textual and acoustic features. The latter is a lightweight model optimized to efficiently recognize clear and intuitive emotions, like ‘anger’ and ‘happiness’, through simple element-wise operations. The final prediction is derived from an average ensemble of logits from both sub-models, ensuring a robust and balanced classification. Evaluated on the Korean AI-Hub and English IEMOCAP datasets, the framework achieves state-of-the-art accuracy of 0.8071 and demonstrates excellent cross-lingual generalization with an accuracy of 0.7823. Empirical results validate the complementary design, confirming that the specialized models synergistically enhance overall performance.

Cite

CITATION STYLE

APA

Yi, M. H., Kwak, K. C., & Shin, J. H. (2025). A Complementary Fusion Framework for Robust Multimodal Emotion Recognition. Electronics (Switzerland), 14(22). https://doi.org/10.3390/electronics14224444

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free