Generating question relevant captions to aid visual question answering

29Citations
Citations of this article
178Readers
Mendeley users who have this article in their library.

Abstract

Visual question answering (VQA) and image captioning require a shared body of general knowledge connecting language and vision. We present a novel approach to improve VQA performance that exploits this connection by jointly generating captions that are targeted to help answer a specific visual question. The model is trained using an existing caption dataset by automatically determining question-relevant captions using an online gradient-based method. Experimental results on the VQA v2 challenge demonstrates that our approach obtains state-of-the-art VQA performance (e.g. 68.4% on the Test-standard set using a single model) by simultaneously generating question-relevant captions.

Cite

CITATION STYLE

APA

Wu, J., Hu, Z., & Mooney, R. J. (2020). Generating question relevant captions to aid visual question answering. In ACL 2019 - 57th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (pp. 3585–3594). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p19-1348

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free