Abstract
AbstractDespite the immense potential of using generative artificial intelligence in architectural design automation, aligning architects' design intents through text-to-image generation remains a challenge, especially in the early stages of architectural design. This work especially focuses on campus buildings and proposes a multi-stage retrieval-augmented diffusion model for fine-grained control of vertical configuration, detailed components (such as windows and doors), and rendering styles in architectural image generation. The proposed framework first generates controlled vertical sketches by integrating style and retrieved structural information with floor numbers. These sketches are then refined into component-controlled sketches via component segmentation and retrieval. Finally, component-controlled sketches are used to generate style-controlled renderings along with text prompts. Experiment results validated the framework's superior performance and controllability in high-quality campus building image generation. We believe that the proposed framework can facilitate the automation of building designs using generative models, enabling robust generation of detailed structural and rendering control.
Author supplied keywords
Cite
CITATION STYLE
Wang, Z., Ren, Y., Jin, H., Feng, J., Du, X., Zhang, Y., & Xie, H. (2026). Controllable generation of building representations: Aligning campus building design intent with multi-stage retrieval-augmented diffusion models. Frontiers of Architectural Research. https://doi.org/10.1016/j.foar.2026.01.018
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.