Abstract
The integration of texts, images, and styles into one single format has posed as a challenge for researchers in text, image, and style synthesis. In this simulation, we present a multimodal structure using Generative Adversarial Networks (GAN) for image synthesis that address the incorporation of textual depiction, reference images, and style into one defined image. In this study we designed a text encoder, a style integration model alongside an image feature extractor to ensure that all images generated meet industry standards of style and quality. Through the studied methods of adversarial loss, text image consistency loss, style matching loss, and several other additional loss functions, we were able to optimize the generation process towards higher precision standards. Results clearly indicate an unpaired multimodal approach combined with our method yielded sharper and more consistent images when validated on a variety of public datasets affording our method an edge over competing theories and methods. The conclusions drawn within this research highlight the gap present in existing literature regarding multimodal image generation while showcasing its wide ranging applications.
Author supplied keywords
Cite
CITATION STYLE
Tan, C., Zhang, W., Qi, Z., Shih, K., Li, X., & Xiang, A. (2025). Generating Multimodal Images with GAN: Integrating Text, Image, and Style. In Proceedings of the 2025 2nd International Conference on Computer and Multimedia Technology, ICCMT 2025 (pp. 16–21). Association for Computing Machinery, Inc. https://doi.org/10.1145/3757749.3757753
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.