Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer for Social Media Popularity Prediction

12Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Social media popularity prediction aims to predict future interaction or attractiveness of new posts. However, in most existing works, there is a notable deficiency in the effective treatment of numerical features. Despite their significant potential to provide ample information, these features are often inadequately processed, leading to insufficiency of information acquirement. In this paper, we introduce a method, named Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer (DFT-MOVLT). To supplement the information in vision-and-language pre-training (VLP), we propose compound text, which is concatenated by numerical data and text. Furthermore, during VLP, a transformer is trained using 3 objectives to ensure thorough feature extraction. Finally, for more generalized prediction, we fine-tune 2 models using different training ways and ensemble them. To evaluate the effectiveness of each mechanism adopted in the proposed method, we conduct an array of ablation experiments. Our team achieve the 3rd place in Social Media Prediction (SMP) Challenge 2023.

Cite

CITATION STYLE

APA

Chen, X., Chen, W., Huang, C., Zhang, Z., Duan, L., & Zhang, Y. (2023). Double-Fine-Tuning Multi-Objective Vision-and-Language Transformer for Social Media Popularity Prediction. In MM 2023 - Proceedings of the 31st ACM International Conference on Multimedia (pp. 9462–9466). Association for Computing Machinery, Inc. https://doi.org/10.1145/3581783.3612845

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free