Abstract
This paper presents a new image captioning system which contains facial expression recognition as a way to provide better emotional and contextual comprehension of the captions generated. A combination of affective cues and visual features is made, which enables semantically full and emotionally conscious descriptions. Experiments were carried out on two created datasets, FlickrFace11k and COCOFace15k, with standard benchmarks such as BLEU, METEOR, ROUGE-L, CIDEr, and SPICE to analyze their effectiveness. The suggested model produced better results in all metrics as compared to baselines, like Show-Attend-Tell and Up-Down, remaining consistently better on all the scores. Remarkably, it has reached gains of 2.5 points on CIDEr and 1.0 on SPICE, which means a closer correlation to the prompt captions made by people. A 5-fold cross-validation confirmed the model’s robustness, with minimal standard deviation across folds (
Author supplied keywords
Cite
CITATION STYLE
Khan, A. S., Khan, A. H., Abbass, M. J., & Shafi, I. (2025). Image Captioning with Object Detection and Facial Expression Recognition for Smart Industry. Bioengineering, 12(12). https://doi.org/10.3390/bioengineering12121325
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.