Image generation based on image description using artificial intelligence

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Purpose – The purpose of this study is to explore and evaluate advanced text-to-image synthesis methods that generate realistic and semantically aligned images from textual descriptions. By leveraging modern deep learning approaches, the research aims to improve image generation quality, diversity and textual coherence using artificial intelligence techniques. Design/methodology/approach – The research focuses on designing, implementing and training four different text-to-image generator architectures based on generative adversarial networks (GANs) and transformer-based text embeddings. Two distinct text-to-image fusion strategies were applied: deep fusion (DF) via affine transformations and a semantic-spatial attention mechanism. The models were trained on three large datasets (CUB-200, MS-COCO and ImageNet), resulting in 12 unique generator-discriminator configurations. Performance was evaluated using the inception score and Fréchet inception distance (FID). Findings – The proposed architecture, combining DF blocks and Semantic-Spatial Aware Convolution Network (SSACN) blocks, achieved competitive results, outperforming several existing models such as AttnGAN, MirrorGAN and DF-GAN in terms of FID. The best-performing model demonstrated its ability to generate diverse and high-quality images that are semantically consistent with the input captions. The use of semantic-spatial fusion further improved the focus and alignment of generated content to the relevant regions described in the text. Originality/value – This work contributes to the field of text-to-image synthesis by introducing and experimentally validating a hybrid fusion approach that integrates global and spatially aware semantic conditioning. The developed models, supported by a systematic evaluation across multiple datasets, demonstrate improved performance over several state-of-the-art solutions, offering a valuable framework for future research in multimodal content generation.

Cite

CITATION STYLE

APA

Šimić, A., & Bagić Babac, M. (2025). Image generation based on image description using artificial intelligence. Applied Computing and Informatics, 1–11. https://doi.org/10.1108/ACI-05-2025-0186

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free