Text Semantics to Image Generation: A Method of Building Facades Design Base on Stable Diffusion Model

42Citations
Citations of this article
32Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Stable Diffusion model has been extensively employed in the study of architectural image generation, but there is still an opportunity to enhance in terms of the controllability of the generated image content. A multi-network combined text-to-building facade image generating method is proposed in this work. We first fine-tuned the Stable Diffusion model on the CMP Facades dataset using the LoRA (Low-Rank Adaptation) approach, then we apply the ControlNet model to further control the output. Finally, we contrasted the facade generating outcomes under various architectural style text contents and control strategies. The results demonstrate that the LoRA training approach significantly decreases the possibility of fine-tuning the Stable Diffusion large model, and the addition of the ControlNet model increases the controllability of the creation of text to building facade images. This provides a foundation for subsequent studies on the generation of architectural images.

Cite

CITATION STYLE

APA

Ma, H., & Zheng, H. (2024). Text Semantics to Image Generation: A Method of Building Facades Design Base on Stable Diffusion Model. In Computational Design and Robotic Fabrication (Vol. Part F2072, pp. 24–34). Springer. https://doi.org/10.1007/978-981-99-8405-3_3

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free