Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense

1Citations
Citations of this article
23Readers
Mendeley users who have this article in their library.

Abstract

Visual commonsense understanding requires Vision Language (VL) models to not only understand image and text but also cross-reference in-between to fully integrate and achieve comprehension of the visual scene described. Recently, various approaches have been developed and have achieved high performance on visual commonsense benchmarks. However, it is unclear whether the models really understand the visual scene and underlying commonsense knowledge due to limited evaluation data resources. To provide an in-depth analysis, we present a Multimodal Evaluation (ME) pipeline to automatically generate question-answer pairs to test models' understanding of the visual scene, text, and related knowledge. We then take a step further to show that training with the ME data boosts model's performance in standard VCR evaluation. Lastly, our in-depth analysis and comparison reveal interesting findings: (1) semantically low-level information can assist learning of high-level information but not the opposite; (2) visual information is generally under utilization compared with text.

Cite

CITATION STYLE

APA

Wang, Z., You, H., He, Y., Li, W., Chang, K. W., & Chang, S. F. (2022). Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022 (pp. 9212–9224). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.emnlp-main.626

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free