Image Captioning with Object Detection and Facial Expression Recognition for Smart Industry

N/ACitations
Citations of this article
10Readers
Mendeley users who have this article in their library.

Abstract

This paper presents a new image captioning system which contains facial expression recognition as a way to provide better emotional and contextual comprehension of the captions generated. A combination of affective cues and visual features is made, which enables semantically full and emotionally conscious descriptions. Experiments were carried out on two created datasets, FlickrFace11k and COCOFace15k, with standard benchmarks such as BLEU, METEOR, ROUGE-L, CIDEr, and SPICE to analyze their effectiveness. The suggested model produced better results in all metrics as compared to baselines, like Show-Attend-Tell and Up-Down, remaining consistently better on all the scores. Remarkably, it has reached gains of 2.5 points on CIDEr and 1.0 on SPICE, which means a closer correlation to the prompt captions made by people. A 5-fold cross-validation confirmed the model’s robustness, with minimal standard deviation across folds (

Cite

CITATION STYLE

APA

Khan, A. S., Khan, A. H., Abbass, M. J., & Shafi, I. (2025). Image Captioning with Object Detection and Facial Expression Recognition for Smart Industry. Bioengineering, 12(12). https://doi.org/10.3390/bioengineering12121325

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free