Multimodal Emotion Recognition via the Fusion of Mamba and Liquid Neural Networks with Cross-Modal Alignment

8Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

This paper proposes a novel multimodal emotion recognition framework, termed Sparse Alignment and Liquid-Mamba (SALM), which effectively integrates the complementary strengths of Mamba networks and Liquid Neural Networks (LNNs). To capture neural dynamics, high-resolution EEG spectrograms are generated via Short-Time Fourier Transform (STFT), while heatmap features from facial images, videos, speech, and text are extracted and aligned through entropy-regularized Sinkhorn and Greenkhorn optimal transport algorithms. These aligned representations are fused to mitigate semantic disparities across modalities. The proposed SALM model leverages sparse alignment for efficient cross-modal mapping and employs the Liquid-Mamba architecture to construct a robust and generalizable classifier. Extensive experiments on benchmark datasets demonstrate that SALM consistently outperforms state-of-the-art methods in both classification accuracy and generalization ability.

Cite

CITATION STYLE

APA

Chen, G., Liao, Y., Zhang, D., Yang, W., Mai, Z., & Xu, C. (2025). Multimodal Emotion Recognition via the Fusion of Mamba and Liquid Neural Networks with Cross-Modal Alignment. Electronics (Switzerland), 14(18). https://doi.org/10.3390/electronics14183638

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free