Abstract
Multimodal sentiment analysis (MSA) seeks to predict subjective human sentiments by utilizing information from multiple modalities. It has been applied in diverse scenarios. Recent studies suggest that MSA benefits from integrating diverse modalities, emphasizing the fusion of multimodal information at different levels and the joint learning of modality-consistent and modality-inconsistent representations. In this work, we put forward a cross-modality dual-attention fusion network (CMDAF) that incorporates a cross-modality dual-attention fusion module. This module facilitates the integration of representations across modalities, enabling effective fusion from shallow to deep levels. A Transformer-based projection module transforms the fused representations into modality-consistent and modality-inconsistent forms, further enhanced through multi-scale contrast learning. Experiments on three MSA datasets (MOSI, MOSEI, and CH-SIMS) demonstrate the superiority of CMDAF, achieving state-of-the-art performance on most metrics.
Author supplied keywords
Cite
CITATION STYLE
Guo, W., Su, K., Jiang, B., Xie, K., & Liu, J. (2024). CMDAF: Cross-Modality Dual-Attention Fusion Network for Multimodal Sentiment Analysis. Applied Sciences (Switzerland), 14(24). https://doi.org/10.3390/app142412025
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.