Abstract
Recent advancements in diffusion models have significantly advanced text-to-image generation, yet global text prompts alone remain insufficient for achieving fine-grained control over individual entities within an image. To address this limitation, we present EliGen, a novel framework for Entity-level controlled image Generation. Firstly, we put forward regional attention, a mechanism for diffusion transformers that requires no additional structures, seamlessly integrating entity prompts and arbitrary-shaped spatial masks. By contributing a high-quality dataset with fine-grained spatial and semantic entity-level annotations, we train EliGen to achieve robust and accurate entity-level manipulation, surpassing existing methods in both spatial precision and image quality. Additionally, we propose an inpainting fusion pipeline, extending EliGen's capabilities to multi-entity image inpainting tasks. We further demonstrate EliGen's flexibility by integrating it with other open-source models such as IP-Adapter, In-Context LoRA and MLLM, unlocking new creative possibilities. The source code, model, and dataset will be published.
Author supplied keywords
Cite
CITATION STYLE
Zhang, H., Duan, Z., Wang, X., Chen, Y., & Zhang, Y. (2025). EliGen: Entity-Level Controlled Image Generation with Regional Attention. In Proceedings of the 7th ACM International Conference on Multimedia in Asia, MMAsia 2025. Association for Computing Machinery, Inc. https://doi.org/10.1145/3743093.3771013
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.