Attention correctness in neural image captioning

N/ACitations
Citations of this article
226Readers
Mendeley users who have this article in their library.

Abstract

Attention mechanisms have recently been introduced in deep learning for various tasks in natural language processing and computer vision. But despite their popularity, the "correctness" of the implicitly-learned attention maps has only been assessed qualitatively by visualization of several examples. In this paper we focus on evaluating and improving the correctness of attention in neural image captioning models. Specifically, we propose a quantitative evaluation metric for the consistency between the generated attention maps and human annotations, using recently released datasets with alignment between regions in images and entities in captions. We then propose novel models with different levels of explicit supervision for learning attention maps during training. The supervision can be strong when alignment between regions and caption entities are available, or weak when only object segments and categories are provided. We show on the popular Flickr30k and COCO datasets that introducing supervision of attention maps during training solidly improves both attention correctness and caption quality, showing the promise of making machine perception more human-like.

Cite

CITATION STYLE

APA

Liu, C., Mao, J., Sha, F., & Yuille, A. (2017). Attention correctness in neural image captioning. In 31st AAAI Conference on Artificial Intelligence, AAAI 2017 (pp. 4176–4182). AAAI press. https://doi.org/10.1609/aaai.v31i1.11197

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free