Abstract
Multi-modal sarcasm detection involves determining whether a given multi-modal input conveys sarcastic intent by analyzing the underlying sentiment. Recently, vision large language models have shown remarkable success on various of multi-modal tasks. Inspired by this, we systematically investigate the impact of vision large language models in zero-shot multi-modal sarcasm detection task. Furthermore, to capture different perspectives of sarcastic expressions, we propose a multi-view agent framework, S3 Agent, designed to enhance zero-shot multi-modal sarcasm detection by leveraging three critical perspectives: superficial expression, semantic information, and sentiment expression. Our experiments on the MMSD2.0 dataset, which involves six models and four prompting strategies, demonstrate that our approach achieves state-of-the-art performance. Our method achieves an average improvement of 13.2% in accuracy. Moreover, we evaluate our method on the text-only sarcasm detection task, where it also surpasses baseline approaches.
Author supplied keywords
Cite
CITATION STYLE
Wang, P., Zhang, Y., Fei, H., Chen, Q., Wang, Y., Si, J., … Qin, L. (2025). S3 Agent: Unlocking the Power of VLLM for Zero-Shot Multi-Modal Sarcasm Detection. ACM Transactions on Multimedia Computing, Communications and Applications, 21(11). https://doi.org/10.1145/3690642
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.