StyleMamba : State Space Model for Efficient Text-driven Image Style Transfer

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Zijia, Liu, Zhi-Song
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911870664835072
author Wang, Zijia
Liu, Zhi-Song
author_facet Wang, Zijia
Liu, Zhi-Song
contents We present StyleMamba, an efficient image style transfer framework that translates text prompts into corresponding visual styles while preserving the content integrity of the original images. Existing text-guided stylization requires hundreds of training iterations and takes a lot of computing resources. To speed up the process, we propose a conditional State Space Model for Efficient Text-driven Image Style Transfer, dubbed StyleMamba, that sequentially aligns the image features to the target text prompts. To enhance the local and global style consistency between text and image, we propose masked and second-order directional losses to optimize the stylization direction to significantly reduce the training iterations by 5 times and the inference time by 3 times. Extensive experiments and qualitative evaluation confirm the robust and superior stylization performance of our methods compared to the existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2405_05027
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StyleMamba : State Space Model for Efficient Text-driven Image Style Transfer
Wang, Zijia
Liu, Zhi-Song
Computer Vision and Pattern Recognition
Artificial Intelligence
We present StyleMamba, an efficient image style transfer framework that translates text prompts into corresponding visual styles while preserving the content integrity of the original images. Existing text-guided stylization requires hundreds of training iterations and takes a lot of computing resources. To speed up the process, we propose a conditional State Space Model for Efficient Text-driven Image Style Transfer, dubbed StyleMamba, that sequentially aligns the image features to the target text prompts. To enhance the local and global style consistency between text and image, we propose masked and second-order directional losses to optimize the stylization direction to significantly reduce the training iterations by 5 times and the inference time by 3 times. Extensive experiments and qualitative evaluation confirm the robust and superior stylization performance of our methods compared to the existing baselines.
title StyleMamba : State Space Model for Efficient Text-driven Image Style Transfer
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2405.05027