Beyond Color and Lines: Zero-Shot Style-Specific Image Variations with Coordinated Semantics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Jinghao, Zhang, Yuhe, Geng, GuoHua, Yang, Liuyuxin, Yan, JiaRui, Cheng, Jingtao, Zhang, YaDong, Li, Kang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929556829503488
author Hu, Jinghao
Zhang, Yuhe
Geng, GuoHua
Yang, Liuyuxin
Yan, JiaRui
Cheng, Jingtao
Zhang, YaDong
Li, Kang
author_facet Hu, Jinghao
Zhang, Yuhe
Geng, GuoHua
Yang, Liuyuxin
Yan, JiaRui
Cheng, Jingtao
Zhang, YaDong
Li, Kang
contents Traditionally, style has been primarily considered in terms of artistic elements such as colors, brushstrokes, and lighting. However, identical semantic subjects, like people, boats, and houses, can vary significantly across different artistic traditions, indicating that style also encompasses the underlying semantics. Therefore, in this study, we propose a zero-shot scheme for image variation with coordinated semantics. Specifically, our scheme transforms the image-to-image problem into an image-to-text-to-image problem. The image-to-text operation employs vision-language models e.g., BLIP) to generate text describing the content of the input image, including the objects and their positions. Subsequently, the input style keyword is elaborated into a detailed description of this style and then merged with the content text using the reasoning capabilities of ChatGPT. Finally, the text-to-image operation utilizes a Diffusion model to generate images based on the text prompt. To enable the Diffusion model to accommodate more styles, we propose a fine-tuning strategy that injects text and style constraints into cross-attention. This ensures that the output image exhibits similar semantics in the desired style. To validate the performance of the proposed scheme, we constructed a benchmark comprising images of various styles and scenes and introduced two novel metrics. Despite its simplicity, our scheme yields highly plausible results in a zero-shot manner, particularly for generating stylized images with high-fidelity semantics.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18537
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Color and Lines: Zero-Shot Style-Specific Image Variations with Coordinated Semantics
Hu, Jinghao
Zhang, Yuhe
Geng, GuoHua
Yang, Liuyuxin
Yan, JiaRui
Cheng, Jingtao
Zhang, YaDong
Li, Kang
Computer Vision and Pattern Recognition
68T07
Traditionally, style has been primarily considered in terms of artistic elements such as colors, brushstrokes, and lighting. However, identical semantic subjects, like people, boats, and houses, can vary significantly across different artistic traditions, indicating that style also encompasses the underlying semantics. Therefore, in this study, we propose a zero-shot scheme for image variation with coordinated semantics. Specifically, our scheme transforms the image-to-image problem into an image-to-text-to-image problem. The image-to-text operation employs vision-language models e.g., BLIP) to generate text describing the content of the input image, including the objects and their positions. Subsequently, the input style keyword is elaborated into a detailed description of this style and then merged with the content text using the reasoning capabilities of ChatGPT. Finally, the text-to-image operation utilizes a Diffusion model to generate images based on the text prompt. To enable the Diffusion model to accommodate more styles, we propose a fine-tuning strategy that injects text and style constraints into cross-attention. This ensures that the output image exhibits similar semantics in the desired style. To validate the performance of the proposed scheme, we constructed a benchmark comprising images of various styles and scenes and introduced two novel metrics. Despite its simplicity, our scheme yields highly plausible results in a zero-shot manner, particularly for generating stylized images with high-fidelity semantics.
title Beyond Color and Lines: Zero-Shot Style-Specific Image Variations with Coordinated Semantics
topic Computer Vision and Pattern Recognition
68T07
url https://arxiv.org/abs/2410.18537