MixSA: Training-free Reference-based Sketch Extraction via Mixture-of-Self-Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Rui, Wu, Xiaojun, He, Shengfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917882688962560
author Yang, Rui
Wu, Xiaojun
He, Shengfeng
author_facet Yang, Rui
Wu, Xiaojun
He, Shengfeng
contents Current sketch extraction methods either require extensive training or fail to capture a wide range of artistic styles, limiting their practical applicability and versatility. We introduce Mixture-of-Self-Attention (MixSA), a training-free sketch extraction method that leverages strong diffusion priors for enhanced sketch perception. At its core, MixSA employs a mixture-of-self-attention technique, which manipulates self-attention layers by substituting the keys and values with those from reference sketches. This allows for the seamless integration of brushstroke elements into initial outline images, offering precise control over texture density and enabling interpolation between styles to create novel, unseen styles. By aligning brushstroke styles with the texture and contours of colored images, particularly in late decoder layers handling local textures, MixSA addresses the common issue of color averaging by adjusting initial outlines. Evaluated with various perceptual metrics, MixSA demonstrates superior performance in sketch quality, flexibility, and applicability. This approach not only overcomes the limitations of existing methods but also empowers users to generate diverse, high-fidelity sketches that more accurately reflect a wide range of artistic expressions.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00816
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MixSA: Training-free Reference-based Sketch Extraction via Mixture-of-Self-Attention
Yang, Rui
Wu, Xiaojun
He, Shengfeng
Computer Vision and Pattern Recognition
Current sketch extraction methods either require extensive training or fail to capture a wide range of artistic styles, limiting their practical applicability and versatility. We introduce Mixture-of-Self-Attention (MixSA), a training-free sketch extraction method that leverages strong diffusion priors for enhanced sketch perception. At its core, MixSA employs a mixture-of-self-attention technique, which manipulates self-attention layers by substituting the keys and values with those from reference sketches. This allows for the seamless integration of brushstroke elements into initial outline images, offering precise control over texture density and enabling interpolation between styles to create novel, unseen styles. By aligning brushstroke styles with the texture and contours of colored images, particularly in late decoder layers handling local textures, MixSA addresses the common issue of color averaging by adjusting initial outlines. Evaluated with various perceptual metrics, MixSA demonstrates superior performance in sketch quality, flexibility, and applicability. This approach not only overcomes the limitations of existing methods but also empowers users to generate diverse, high-fidelity sketches that more accurately reflect a wide range of artistic expressions.
title MixSA: Training-free Reference-based Sketch Extraction via Mixture-of-Self-Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.00816