Abstract
This paper proposes a multi-scale adaptive fusion sarcasm detection model (MSAF-SDM) to address the challenges of information complexity and insufficient inter-modal collaboration in multimodal sarcasm detection. The model integrates multi-level features from text, audio, and video modalities, leveraging a dynamic attention mechanism and an adaptive weight allocation strategy to capture sarcasm-related cues across modalities. To enhance feature extraction capabilities, the text modality employs a dual multi-scale dilated window attention mechanism, the audio modality utilizes multi-scale temporal convolution, and the video modality incorporates multi-scale spatiotemporal convolution reinforced by auxiliary modal features. Experimental results demonstrate that MSAF-SDM achieves an accuracy of 89.04% and an F1-score of 87.68% on public datasets, significantly outperforming existing state-of-the-art models. Ablation studies further validate the effectiveness of the multimodal feature extraction and adaptive fusion mechanisms. This research provides a novel approach for tackling multimodal sarcasm detection tasks.
Author supplied keywords
Cite
CITATION STYLE
Wu, H., & Zang, Y. (2025). A multi-scale adaptive fusion model for multimodal sarcasm detection. Discover Computing, 28(1). https://doi.org/10.1007/s10791-025-09730-y
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.