Abstract
To address issues such as missing modalities and information conflicts in multimodal emotion research, this study proposes a Dynamically Adaptive Multimodal Fusion Transformer (DAMFT). DAMFT effectively integrates the complementary advantages of visual, audio, and text data through a multimodal fusion mechanism. It uses linear transformations and attention mechanisms to unify and align multimodal emotions, and assesses emotional intensity by comparing the absolute values of emotions, thereby achieving emotion intensity-driven feature interaction. Based on this, it guides auxiliary modalities to generate hypermodal representations using the primary modality, realizing adaptively guided cross-modal feature information complementarity, and solving the problems of information loss and interference in multimodal sentiment analysis. Finally, extensive experiments demonstrate the excellent performance of the proposed model on multiple datasets, and ablation experiments prove the contribution of each module in the model, achieving more accurate and robust sentiment analysis performance.
Author supplied keywords
Cite
CITATION STYLE
Wu, W., Hu, Q., Shen, X., & Zhang, Z. (2025). Dynamically Adaptive Multimodal Fusion Transformer for Multimodal Sentiment Analysis. IEEE Access, 13, 193754–193764. https://doi.org/10.1109/ACCESS.2025.3628494
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.