One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Han, Li, Junyao, Zhao, Kangbo, Zhang, Sen, Shi, Yukai, Lin, Liang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918062837465088
author Wu, Han
Li, Junyao
Zhao, Kangbo
Zhang, Sen
Shi, Yukai
Lin, Liang
author_facet Wu, Han
Li, Junyao
Zhao, Kangbo
Zhang, Sen
Shi, Yukai
Lin, Liang
contents Face sketch synthesis is a technique aimed at converting face photos into sketches. Existing face sketch synthesis research mainly relies on training with numerous photo-sketch sample pairs from existing datasets. However, these large-scale discriminative learning methods will have to face problems such as data scarcity and high human labor costs. Once the training data becomes scarce, their generative performance significantly degrades. In this paper, we propose a one-shot face sketch synthesis method based on diffusion models. We optimize text instructions on a diffusion model using face photo-sketch image pairs. Then, the instructions derived through gradient-based optimization are used for inference. To simulate real-world scenarios more accurately and evaluate method effectiveness more comprehensively, we introduce a new benchmark named One-shot Face Sketch Dataset (OS-Sketch). The benchmark consists of 400 pairs of face photo-sketch images, including sketches with different styles and photos with different backgrounds, ages, sexes, expressions, illumination, etc. For a solid out-of-distribution evaluation, we select only one pair of images for training at each time, with the rest used for inference. Extensive experiments demonstrate that the proposed method can convert various photos into realistic and highly consistent sketches in a one-shot context. Compared to other methods, our approach offers greater convenience and broader applicability. The dataset will be available at: https://github.com/HanWu3125/OS-Sketch
format Preprint
id arxiv_https___arxiv_org_abs_2506_15312
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning
Wu, Han
Li, Junyao
Zhao, Kangbo
Zhang, Sen
Shi, Yukai
Lin, Liang
Graphics
Cryptography and Security
Computer Vision and Pattern Recognition
Computers and Society
Face sketch synthesis is a technique aimed at converting face photos into sketches. Existing face sketch synthesis research mainly relies on training with numerous photo-sketch sample pairs from existing datasets. However, these large-scale discriminative learning methods will have to face problems such as data scarcity and high human labor costs. Once the training data becomes scarce, their generative performance significantly degrades. In this paper, we propose a one-shot face sketch synthesis method based on diffusion models. We optimize text instructions on a diffusion model using face photo-sketch image pairs. Then, the instructions derived through gradient-based optimization are used for inference. To simulate real-world scenarios more accurately and evaluate method effectiveness more comprehensively, we introduce a new benchmark named One-shot Face Sketch Dataset (OS-Sketch). The benchmark consists of 400 pairs of face photo-sketch images, including sketches with different styles and photos with different backgrounds, ages, sexes, expressions, illumination, etc. For a solid out-of-distribution evaluation, we select only one pair of images for training at each time, with the rest used for inference. Extensive experiments demonstrate that the proposed method can convert various photos into realistic and highly consistent sketches in a one-shot context. Compared to other methods, our approach offers greater convenience and broader applicability. The dataset will be available at: https://github.com/HanWu3125/OS-Sketch
title One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning
topic Graphics
Cryptography and Security
Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2506.15312