Abstract
With the increasing prevalence of multimodal user-generated content on social media, Multimodal Aspect-Based Sentiment Analysis (MABSA) has garnered significant attention in recent years. MABSA aims to classify the sentiment polarity of aspects mentioned in textual content by leveraging both textual and visual modalities. Previous studies have primarily focused on using transformer-based models to fuse information from different modalities. However, the presence of irrelevant information in images and text related to specific aspects can negatively impact sentiment analysis results. Many studies still face challenges in handling the noise introduced during the multimodal fusion process. To address this issue, we propose a novel Position-Perceptive Multi-hop Fusion Network (PPMFN) in this paper. Our method includes a position-perceptive module and an aspect-guided multi-hop interactive attention fusion module, aimed at enhancing modal position perception and aspect-guided multi-hop interactions to filter out irrelevant noise information. Typically, parts closer to the target are more relevant. The position-perceptive module is used to capture positional features from both images and text. The aspect-guided fusion module leverages a multi-hop interactive attention network to merge multimodal features while eliminating irrelevant information from both images and text. Extensive experimental results demonstrate that our framework consistently surpasses strong baseline models across two public datasets.
Author supplied keywords
Cite
CITATION STYLE
Fan, H., & Chen, J. (2024). Position Perceptive Multi-Hop Fusion Network for Multimodal Aspect-Based Sentiment Analysis. IEEE Access, 12, 90586–90595. https://doi.org/10.1109/ACCESS.2024.3404261
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.