Image caption generator using deep learning

17Citations
Citations of this article
24Readers
Mendeley users who have this article in their library.

Abstract

Computer Vision and Natural Language Processing in artificial intelligence is used for automatically describing the content of an image. In order to describe the image a well-formed English phrases is needed. Automatically describing image content is very much helpful to the visually impaired people to understand the problem better. The paper is intended to identify objects and inform people through audio and text messages. It recognizes image and converts to audio using GTTS and converts to text using LSTM. Initially, the input image is converted to a grayscale image that is processed through the Convolution Neural Network (CNN) to correctly identify the objects. Objects in the image are correctly identified using OpenCV, which is then converted to audio and text messages. The proposed method for blind people is designed to expand to people with vision loss in order to achieve their full potential.

Cite

CITATION STYLE

APA

Krishnakumar, B., Kousalya, K., Gokul, S., Karthikeyan, R., & Kaviyarasu, D. (2020). Image caption generator using deep learning. International Journal of Advanced Science and Technology, 29(3 Special Issue), 975–980. https://doi.org/10.55041/ijsrem31987

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free