A comprehensive review and open challenges on visual question answering models

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

Users are now able to actively interact with images and pose different questions based on images, thanks to recent developments in artificial intelligence. In turn, a response in a natural language answer is expected. The study discusses a variety of datasets that can be used to examine applications for visual question-answering (VQA), as well as their advantages and disadvantages. Four different forms of VQA models - simple joint embedding-based models, attention-based models, knowledge-incorporated models, and domain-specific VQA models - are in-depth examined in this article. We also critically assess the drawbacks and future possibilities of all current state-of-the-art (SoTa), end-to-end VQA models. Finally, we present the directions and guidelines for further development of the VQA models.

Cite

CITATION STYLE

APA

Koshti, D., Gupta, A., Kalla, M., & Sharma, A. (2023). A comprehensive review and open challenges on visual question answering models. Aibi, Revista de Investigacion Administracion e Ingenierias, 11(3), 126–142. https://doi.org/10.15649/2346030X.3370

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free