All You May Need for VQA are Image Captions

N/ACitations
Citations of this article
80Readers
Mendeley users who have this article in their library.

Abstract

Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. In this paper, we propose a method that automatically derives VQA examples at volume, by leveraging the abundance of existing image-caption annotations combined with neural models for textual question generation. We show that the resulting data is of high-quality. VQA models trained on our data improve state-of-the-art zero-shot accuracy by double digits and achieve a level of robustness that lacks in the same model trained on human-annotated VQA data.

Cite

CITATION STYLE

APA

Changpinyo, S., Kukliansky, D., Szpektor, I., Chen, X., Ding, N., & Soricut, R. (2022). All You May Need for VQA are Image Captions. In NAACL 2022 - 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference (pp. 1947–1963). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.naacl-main.142

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free