Beyond Spurious Cues: Adaptive Multi-Modal Fusion via Mixture-of-Experts for Robust Sarcasm Detection

N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Sarcasm is a complex emotional expression often marked by semantic contrast and incongruity between textual and visual modalities. In recent years, multi-modal sarcasm detection (MMSD) has emerged as a vital task in affective computing. However, existing models frequently rely on superficial spurious cues—such as emojis or hashtags—during training and inference, limiting their ability to capture deeper semantic inconsistencies and undermining generalization to real-world scenarios. To tackle these challenges, we propose Multi-Modal Mixture-of-Experts (MM-MoE), a novel framework that integrates diverse expert modules through a global dynamic gating mechanism for adaptive cross-modal interaction and selective semantic fusion. This architecture allows for the model to better capture modality-level incongruity. Furthermore, we introduce MMSD3.0 and MMSD4.0, two cross-dataset evaluation benchmarks derived from two open source benchmark datasets, MMSD and MMSD2.0, to assess model robustness under varying distributions of spurious cues. Extensive experiments demonstrate that MM-MoE achieves strong performance and generalization ability, consistently outperforming state-of-the-art baselines when encountering superficial spurious correlations.

Cite

CITATION STYLE

APA

Zhao, G., Zhao, Y., Yin, X., Lin, L., & Zhu, J. (2025). Beyond Spurious Cues: Adaptive Multi-Modal Fusion via Mixture-of-Experts for Robust Sarcasm Detection. Mathematics, 13(20). https://doi.org/10.3390/math13203250

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free