Multilingual visual question answering for visually impaired people

2Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Visual question answering (VQA) aims to answer questions for a given image. The applications of VQA systems are well explored in education, e-commerce, and interactive exhibits. It also enhances accessibility for the visually impaired (VI). Several VQA systems exist in English for various applications. However, VQA developed for VI people is limited, and such VQA in low-resource languages, specifically Hindi and Bengali, does not exist. This article introduces two such datasets in Bengali and Hindi. The datasets are machine-translated from the popular VQA-VI dataset VizWiz, and curated by native speakers. The datasets consist of approximately 20K image-question pairs along with 10 different answers. We also report benchmark results using state-of-the-art VQA methods and explore different pre-trained embeddings. We achieve a maximum answer type prediction accuracy and answer accuracy of 68.00%/20.35% (Bengali) and 67.09%/23.06% (Hindi). The low accuracy using recent state-of-the-art methods is evidence of the complexity of the datasets. We hope the datasets will attract researchers and create a baseline for VQA for VI people in resource-constrained Indic languages. The code and the datasets are available in url. The URL (will be updated) when published.

Cite

CITATION STYLE

APA

Pal, R., Kar, S., Prasad, D. K., & Sekh, A. A. (2025). Multilingual visual question answering for visually impaired people. Discover Artificial Intelligence, 5(1). https://doi.org/10.1007/s44163-025-00482-8

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free