ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors

48Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Recently, the multimedia community has witnessed the rise of diffusion models trained on large-scale multi-modal data for visual content creation, particularly in the field of text-to-image generation. In this paper, we propose a new task for "stylizing'' text-to-image models, namely text-driven stylized image generation, that further enhances editability in content creation. Given input text prompt and style image, this task aims to produce stylized images which are both semantically relevant to input text prompt and meanwhile aligned with the style image in style. To achieve this, we present a new diffusion model (ControlStyle) via upgrading a pre-trained text-to-image model with a trainable modulation network enabling more conditions of text prompts and style images. Moreover, diffusion style and content regularizations are simultaneously introduced to facilitate the learning of this modulation network with these diffusion priors, pursuing high-quality stylized text-to-image generation. Extensive experiments demonstrate the effectiveness of our ControlStyle in producing more visually pleasing and artistic results, surpassing a simple combination of text-to-image model and conventional style transfer techniques.

Cite

CITATION STYLE

APA

Chen, J., Pan, Y., Yao, T., & Mei, T. (2023). ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors. In MM 2023 - Proceedings of the 31st ACM International Conference on Multimedia (pp. 7540–7548). Association for Computing Machinery, Inc. https://doi.org/10.1145/3581783.3612524

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free