Abstract
This paper presents a novel dual-stream framework for multimodal emotion recognition, engineered to address the varying complexity inherent in emotional expressions. The proposed architecture uniquely integrates a graph embedding as an auxiliary modality to explicitly model temporal correlations between utterances, and processes features through two complementary sub-models operating in parallel: a Cross-Attention Transformer Mixture-of-Experts (MoE) model and a Sum-Product Linear MoE model. The former deciphers nuanced and ambiguous emotions, such as ‘fear’ and ‘sadness’, by leveraging a deep cross-attention mechanism to model intricate, bidirectional dependencies between textual and acoustic features. The latter is a lightweight model optimized to efficiently recognize clear and intuitive emotions, like ‘anger’ and ‘happiness’, through simple element-wise operations. The final prediction is derived from an average ensemble of logits from both sub-models, ensuring a robust and balanced classification. Evaluated on the Korean AI-Hub and English IEMOCAP datasets, the framework achieves state-of-the-art accuracy of 0.8071 and demonstrates excellent cross-lingual generalization with an accuracy of 0.7823. Empirical results validate the complementary design, confirming that the specialized models synergistically enhance overall performance.
Author supplied keywords
Cite
CITATION STYLE
Yi, M. H., Kwak, K. C., & Shin, J. H. (2025). A Complementary Fusion Framework for Robust Multimodal Emotion Recognition. Electronics (Switzerland), 14(22). https://doi.org/10.3390/electronics14224444
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.