DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Ying, Zhuang, Chenyi, Gao, Pan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909354559537152
author Hu, Ying
Zhuang, Chenyi
Gao, Pan
author_facet Hu, Ying
Zhuang, Chenyi
Gao, Pan
contents Style transfer aims to fuse the artistic representation of a style image with the structural information of a content image. Existing methods train specific networks or utilize pre-trained models to learn content and style features. However, they rely solely on textual or spatial representations that are inadequate to achieve the balance between content and style. In this work, we propose a novel and training-free approach for style transfer, combining textual embedding with spatial features and separating the injection of content or style. Specifically, we adopt the BLIP-2 encoder to extract the textual representation of the style image. We utilize the DDIM inversion technique to extract intermediate embeddings in content and style branches as spatial features. Finally, we harness the step-by-step property of diffusion models by separating the injection of content and style in the target branch, which improves the balance between content preservation and style fusion. Various experiments have demonstrated the effectiveness and robustness of our proposed DiffeseST for achieving balanced and controllable style transfer results, as well as the potential to extend to other tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15007
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
Hu, Ying
Zhuang, Chenyi
Gao, Pan
Computer Vision and Pattern Recognition
Multimedia
Style transfer aims to fuse the artistic representation of a style image with the structural information of a content image. Existing methods train specific networks or utilize pre-trained models to learn content and style features. However, they rely solely on textual or spatial representations that are inadequate to achieve the balance between content and style. In this work, we propose a novel and training-free approach for style transfer, combining textual embedding with spatial features and separating the injection of content or style. Specifically, we adopt the BLIP-2 encoder to extract the textual representation of the style image. We utilize the DDIM inversion technique to extract intermediate embeddings in content and style branches as spatial features. Finally, we harness the step-by-step property of diffusion models by separating the injection of content and style in the target branch, which improves the balance between content preservation and style fusion. Various experiments have demonstrated the effectiveness and robustness of our proposed DiffeseST for achieving balanced and controllable style transfer results, as well as the potential to extend to other tasks.
title DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2410.15007