Improving Few-Shot Image Classification Using Machine- and User-Generated Natural Language Descriptions

N/ACitations
Citations of this article
35Readers
Mendeley users who have this article in their library.

Abstract

Humans can obtain the knowledge of novel visual concepts from language descriptions, and we thus use the few-shot image classification task to investigate whether a machine learning model can have this capability. Our proposed model, LIDE (Learning from Image and DEscription), has a text decoder to generate the descriptions and a text encoder to obtain the text representations of machine- or user-generated descriptions. We confirmed that LIDE with machine-generated descriptions outperformed baseline models. Moreover, the performance was improved further with high-quality usergenerated descriptions. The generated descriptions can be viewed as the explanations of the model's predictions, and we observed that such explanations were consistent with prediction results. We also investigated why the language description improved the few-shot image classification performance by comparing the image representations and the text representations in the feature spaces.

Cite

CITATION STYLE

APA

Nishida, K., Nishida, K., & Nishioka, S. (2022). Improving Few-Shot Image Classification Using Machine- and User-Generated Natural Language Descriptions. In Findings of the Association for Computational Linguistics: NAACL 2022 - Findings (pp. 1421–1430). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.findings-naacl.106

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free