Bridging Text and Image for Artist Style Transfer via Contrastive Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zhi-Song, Wang, Li-Wen, Xiao, Jun, Kalogeiton, Vicky
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910730573316096
author Liu, Zhi-Song
Wang, Li-Wen
Xiao, Jun
Kalogeiton, Vicky
author_facet Liu, Zhi-Song
Wang, Li-Wen
Xiao, Jun
Kalogeiton, Vicky
contents Image style transfer has attracted widespread attention in the past few years. Despite its remarkable results, it requires additional style images available as references, making it less flexible and inconvenient. Using text is the most natural way to describe the style. More importantly, text can describe implicit abstract styles, like styles of specific artists or art movements. In this paper, we propose a Contrastive Learning for Artistic Style Transfer (CLAST) that leverages advanced image-text encoders to control arbitrary style transfer. We introduce a supervised contrastive training strategy to effectively extract style descriptions from the image-text model (i.e., CLIP), which aligns stylization with the text description. To this end, we also propose a novel and efficient adaLN based state space models that explore style-content fusion. Finally, we achieve a text-driven image style transfer. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods in artistic style transfer. More importantly, it does not require online fine-tuning and can render a 512x512 image in 0.03s.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09566
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bridging Text and Image for Artist Style Transfer via Contrastive Learning
Liu, Zhi-Song
Wang, Li-Wen
Xiao, Jun
Kalogeiton, Vicky
Computer Vision and Pattern Recognition
Human-Computer Interaction
Image style transfer has attracted widespread attention in the past few years. Despite its remarkable results, it requires additional style images available as references, making it less flexible and inconvenient. Using text is the most natural way to describe the style. More importantly, text can describe implicit abstract styles, like styles of specific artists or art movements. In this paper, we propose a Contrastive Learning for Artistic Style Transfer (CLAST) that leverages advanced image-text encoders to control arbitrary style transfer. We introduce a supervised contrastive training strategy to effectively extract style descriptions from the image-text model (i.e., CLIP), which aligns stylization with the text description. To this end, we also propose a novel and efficient adaLN based state space models that explore style-content fusion. Finally, we achieve a text-driven image style transfer. Extensive experiments demonstrate that our approach outperforms the state-of-the-art methods in artistic style transfer. More importantly, it does not require online fine-tuning and can render a 512x512 image in 0.03s.
title Bridging Text and Image for Artist Style Transfer via Contrastive Learning
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2410.09566