A multi-scale adaptive fusion model for multimodal sarcasm detection

4Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This paper proposes a multi-scale adaptive fusion sarcasm detection model (MSAF-SDM) to address the challenges of information complexity and insufficient inter-modal collaboration in multimodal sarcasm detection. The model integrates multi-level features from text, audio, and video modalities, leveraging a dynamic attention mechanism and an adaptive weight allocation strategy to capture sarcasm-related cues across modalities. To enhance feature extraction capabilities, the text modality employs a dual multi-scale dilated window attention mechanism, the audio modality utilizes multi-scale temporal convolution, and the video modality incorporates multi-scale spatiotemporal convolution reinforced by auxiliary modal features. Experimental results demonstrate that MSAF-SDM achieves an accuracy of 89.04% and an F1-score of 87.68% on public datasets, significantly outperforming existing state-of-the-art models. Ablation studies further validate the effectiveness of the multimodal feature extraction and adaptive fusion mechanisms. This research provides a novel approach for tackling multimodal sarcasm detection tasks.

Cite

CITATION STYLE

APA

Wu, H., & Zang, Y. (2025). A multi-scale adaptive fusion model for multimodal sarcasm detection. Discover Computing, 28(1). https://doi.org/10.1007/s10791-025-09730-y

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free