Image captioning algorithm based on multi-branch CNN and Bi-LSTM

6Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

The development of deep learning and neural networks has brought broad prospects to computer vision and natural language processing. The image captioning task combines cutting-edge methods in two fields. By building an end-to-end encoder-decoder model, its description performance can be greatly improved. In this paper, the multi-branch deep convolutional neural network is used as the encoder to extract image features, and the recurrent neural network is used to generate descriptive text that matches the input image. We conducted experiments on Flickr8k, Flickr30k and MSCOCO datasets. According to the analysis of the experimental results on evaluation metrics, the model proposed in this paper can effectively achieve image caption, and its performance is better than classic image captioning models such as neural image annotation models.

Cite

CITATION STYLE

APA

He, S., Lu, Y., & Chen, S. (2021). Image captioning algorithm based on multi-branch CNN and Bi-LSTM. IEICE Transactions on Information and Systems, E104D(7), 941–947. https://doi.org/10.1587/transinf.2020EDP7227

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free