Abstract
Purpose – The purpose of this study is to explore and evaluate advanced text-to-image synthesis methods that generate realistic and semantically aligned images from textual descriptions. By leveraging modern deep learning approaches, the research aims to improve image generation quality, diversity and textual coherence using artificial intelligence techniques. Design/methodology/approach – The research focuses on designing, implementing and training four different text-to-image generator architectures based on generative adversarial networks (GANs) and transformer-based text embeddings. Two distinct text-to-image fusion strategies were applied: deep fusion (DF) via affine transformations and a semantic-spatial attention mechanism. The models were trained on three large datasets (CUB-200, MS-COCO and ImageNet), resulting in 12 unique generator-discriminator configurations. Performance was evaluated using the inception score and Fréchet inception distance (FID). Findings – The proposed architecture, combining DF blocks and Semantic-Spatial Aware Convolution Network (SSACN) blocks, achieved competitive results, outperforming several existing models such as AttnGAN, MirrorGAN and DF-GAN in terms of FID. The best-performing model demonstrated its ability to generate diverse and high-quality images that are semantically consistent with the input captions. The use of semantic-spatial fusion further improved the focus and alignment of generated content to the relevant regions described in the text. Originality/value – This work contributes to the field of text-to-image synthesis by introducing and experimentally validating a hybrid fusion approach that integrates global and spatially aware semantic conditioning. The developed models, supported by a systematic evaluation across multiple datasets, demonstrate improved performance over several state-of-the-art solutions, offering a valuable framework for future research in multimodal content generation.
Author supplied keywords
Cite
CITATION STYLE
Šimić, A., & Bagić Babac, M. (2025). Image generation based on image description using artificial intelligence. Applied Computing and Informatics, 1–11. https://doi.org/10.1108/ACI-05-2025-0186
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.