V-Shuffle: Zero-Shot Style Transfer via Value Shuffle
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917069740572672 |
|---|---|
| author | Tang, Haojun Lin, Qiwei Xu, Tongda Huang, Lida Wang, Yan |
| author_facet | Tang, Haojun Lin, Qiwei Xu, Tongda Huang, Lida Wang, Yan |
| contents | Attention injection-based style transfer has achieved remarkable progress in recent years. However, existing methods often suffer from content leakage, where the undesired semantic content of the style image mistakenly appears in the stylized output. In this paper, we propose V-Shuffle, a zero-shot style transfer method that leverages multiple style images from the same style domain to effectively navigate the trade-off between content preservation and style fidelity. V-Shuffle implicitly disrupts the semantic content of the style images by shuffling the value features within the self-attention layers of the diffusion model, thereby preserving low-level style representations. We further introduce a Hybrid Style Regularization that complements these low-level representations with high-level style textures to enhance style fidelity. Empirical results demonstrate that V-Shuffle achieves excellent performance when utilizing multiple style images. Moreover, when applied to a single style image, V-Shuffle outperforms previous state-of-the-art methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_06365 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | V-Shuffle: Zero-Shot Style Transfer via Value Shuffle Tang, Haojun Lin, Qiwei Xu, Tongda Huang, Lida Wang, Yan Computer Vision and Pattern Recognition Attention injection-based style transfer has achieved remarkable progress in recent years. However, existing methods often suffer from content leakage, where the undesired semantic content of the style image mistakenly appears in the stylized output. In this paper, we propose V-Shuffle, a zero-shot style transfer method that leverages multiple style images from the same style domain to effectively navigate the trade-off between content preservation and style fidelity. V-Shuffle implicitly disrupts the semantic content of the style images by shuffling the value features within the self-attention layers of the diffusion model, thereby preserving low-level style representations. We further introduce a Hybrid Style Regularization that complements these low-level representations with high-level style textures to enhance style fidelity. Empirical results demonstrate that V-Shuffle achieves excellent performance when utilizing multiple style images. Moreover, when applied to a single style image, V-Shuffle outperforms previous state-of-the-art methods. |
| title | V-Shuffle: Zero-Shot Style Transfer via Value Shuffle |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.06365 |