Abstract
Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. In this paper, we propose a method that automatically derives VQA examples at volume, by leveraging the abundance of existing image-caption annotations combined with neural models for textual question generation. We show that the resulting data is of high-quality. VQA models trained on our data improve state-of-the-art zero-shot accuracy by double digits and achieve a level of robustness that lacks in the same model trained on human-annotated VQA data.
Cite
CITATION STYLE
Changpinyo, S., Kukliansky, D., Szpektor, I., Chen, X., Ding, N., & Soricut, R. (2022). All You May Need for VQA are Image Captions. In NAACL 2022 - 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference (pp. 1947–1963). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.naacl-main.142
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.