CSGO: Content-Style Composition in Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xing, Peng, Wang, Haofan, Sun, Yanpeng, Wang, Qixun, Bai, Xu, Ai, Hao, Huang, Renyuan, Li, Zechao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914935804526592
author Xing, Peng
Wang, Haofan
Sun, Yanpeng
Wang, Qixun
Bai, Xu
Ai, Hao
Huang, Renyuan
Li, Zechao
author_facet Xing, Peng
Wang, Haofan
Sun, Yanpeng
Wang, Qixun
Bai, Xu
Ai, Hao
Huang, Renyuan
Li, Zechao
contents The diffusion model has shown exceptional capabilities in controlled image generation, which has further fueled interest in image style transfer. Existing works mainly focus on training free-based methods (e.g., image inversion) due to the scarcity of specific data. In this study, we present a data construction pipeline for content-style-stylized image triplets that generates and automatically cleanses stylized data triplets. Based on this pipeline, we construct a dataset IMAGStyle, the first large-scale style transfer dataset containing 210k image triplets, available for the community to explore and research. Equipped with IMAGStyle, we propose CSGO, a style transfer model based on end-to-end training, which explicitly decouples content and style features employing independent feature injection. The unified CSGO implements image-driven style transfer, text-driven stylized synthesis, and text editing-driven stylized synthesis. Extensive experiments demonstrate the effectiveness of our approach in enhancing style control capabilities in image generation. Additional visualization and access to the source code can be located on the project page: \url{https://csgo-gen.github.io/}.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16766
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CSGO: Content-Style Composition in Text-to-Image Generation
Xing, Peng
Wang, Haofan
Sun, Yanpeng
Wang, Qixun
Bai, Xu
Ai, Hao
Huang, Renyuan
Li, Zechao
Computer Vision and Pattern Recognition
The diffusion model has shown exceptional capabilities in controlled image generation, which has further fueled interest in image style transfer. Existing works mainly focus on training free-based methods (e.g., image inversion) due to the scarcity of specific data. In this study, we present a data construction pipeline for content-style-stylized image triplets that generates and automatically cleanses stylized data triplets. Based on this pipeline, we construct a dataset IMAGStyle, the first large-scale style transfer dataset containing 210k image triplets, available for the community to explore and research. Equipped with IMAGStyle, we propose CSGO, a style transfer model based on end-to-end training, which explicitly decouples content and style features employing independent feature injection. The unified CSGO implements image-driven style transfer, text-driven stylized synthesis, and text editing-driven stylized synthesis. Extensive experiments demonstrate the effectiveness of our approach in enhancing style control capabilities in image generation. Additional visualization and access to the source code can be located on the project page: \url{https://csgo-gen.github.io/}.
title CSGO: Content-Style Composition in Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.16766