Order-free RNN with visual attention for multi-label classification

N/ACitations
Citations of this article
153Readers
Mendeley users who have this article in their library.

Abstract

We propose a recurrent neural network (RNN) based model for image multi-label classification. Our model uniquely integrates and learning of visual attention and Long Short Term Memory (LSTM) layers, which jointly learns the labels of interest and their co-occurrences, while the associated image regions are visually attended. Different from existing approaches utilize either model in their network architectures, training of our model does not require pre-defined label orders. Moreover, a robust inference process is introduced so that prediction errors would not propagate and thus affect the performance. Our experiments on NUS-WISE and MS-COCO datasets confirm the design of our network and its effectiveness in solving multi-label classification problems.

Cite

CITATION STYLE

APA

Chen, S. F., Chen, Y. C., Yeh, C. K., & Wang, Y. C. F. (2018). Order-free RNN with visual attention for multi-label classification. In 32nd AAAI Conference on Artificial Intelligence, AAAI 2018 (pp. 6714–6721). AAAI press. https://doi.org/10.1609/aaai.v32i1.12230

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free